Natural speech re-synthesis from direct cortical recordings using a pre-trained encoder-decoder framework
Li, J.; Guo, C.; Chang, E. F.; Li, Y.
Show abstract
Reconstructing perceived speech stimuli from neural recordings is not only advancing the understanding of the neural coding underlying speech processing but also an important building block for brain-computer interfaces and neuroprosthetics. However, previous attempts to directly re-synthesize speech from neural decoding suffer from low re-synthesis quality. With the limited neural data and complex speech representation space, it is hard to build decoding model that directly map neural signal into high-fidelity speech. In this work, we proposed a pre-trained encoder-decoder framework to address these problems. We recorded high-density electrocorticography (ECoG) signals when participants listening to natural speech. We built a pre-trained speech re-synthesizing network that consists of a context-dependent speech encoding network and a generative adversarial network (GAN) for high-fidelity speech synthesis. This model was pre-trained on a large naturalistic speech corpus and can extract critical features for speech re-synthesize. We then built a light-weight neural decoding network that mapped the ECoG signal into the latent space of the pre-trained network, and used the GAN decoder to synthesize natural speech. Using only 20 minutes of intracranial neural data, our neural-driven speech re-synthesis model demonstrated promising performance, with phoneme error rate (PER) at 28.6%, and human listeners were able to recognize 71.6% of the words in the re-synthesized speech. This work demonstrates the feasibility of using pre-trained self-supervised model and feature alignment to build efficient neural-to-speech decoding model.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Direct Speech Reconstruction from Sensorimotor Brain Activity with Optimized Deep Learning Models 97%
- Linear versus deep learning methods for noisy speech separation for EEG-informed attention decoding 96%
- Speech decoding from a small set of spatially segregated minimally invasive intracranial EEG electrodes with a compact and interpretable neural network 95%
Similar papers in this journal
- Feasibility of decoding covert speech in ECoG with aTransformer trained on overt speech 97%
- Online speech synthesis using a chronically implanted brain-computer interface in an individual with ALS 96%
- Bridging Auditory Perception and Natural Language Processing with Semantically informed Deep Neural Networks 95%
Similar papers in this journal
- 3 Directional Inception-ResUNet: deep spatial feature learning for multichannel singing voice separation with distortion 94%
- Diffusion model-based image generation from rat brain activity 93%
- Long-term performance assessment of fully automatic biomedical glottis segmentation at the point of care 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.