Speech Synthesis from Electrocorticogram During Imagined Speech Using a Transformer-Based Decoder and Pretrained Vocoder
Komeiji, S.; Shigemi, K.; Mitsuhashi, T.; Iimura, Y.; Suzuki, H.; Sugano, H.; Shinoda, K.; Yatabe, K.; Tanaka, T.
Show abstract
This study presents a method for synthesizing speech from Electrocorticogram signals recorded during imagined speech. To address the limitations posed by the available training data, we employed a Transformer-based decoder to generate log-mel spectrograms, which were then converted into high-quality audio using a pre-trained neural vocoder, Parallel WaveGAN. In experiments involving electrocorticography (ECoG) recordings from 13 participants, the synthesized speech achieved Pearson correlation coefficients ranging from 0.85 to 0.95. These results demonstrate the effectiveness of the Transformer decoder in reconstructing accurate spectrograms from ECoG signals, even under data-constrained conditions.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Comparing MEG and EEG measurement set-ups for a brain--computer interface based on selective auditory attention 95%
- 3 Directional Inception-ResUNet: deep spatial feature learning for multichannel singing voice separation with distortion 95%
- Improving classification and reconstruction of imagined images from EEG signals 95%
Similar papers in this journal
- Feasibility of decoding covert speech in ECoG with aTransformer trained on overt speech 99%
- Online speech synthesis using a chronically implanted brain-computer interface in an individual with ALS 96%
- Deep Learning Restores Speech Intelligibility in Multi-Talker Interference for Cochlear Implant Users 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.