Back

Speech Synthesis from Electrocorticogram During Imagined Speech Using a Transformer-Based Decoder and Pretrained Vocoder

Komeiji, S.; Shigemi, K.; Mitsuhashi, T.; Iimura, Y.; Suzuki, H.; Sugano, H.; Shinoda, K.; Yatabe, K.; Tanaka, T.

2024-08-22 neuroscience
10.1101/2024.08.21.608927 bioRxiv
Show abstract

This study presents a method for synthesizing speech from Electrocorticogram signals recorded during imagined speech. To address the limitations posed by the available training data, we employed a Transformer-based decoder to generate log-mel spectrograms, which were then converted into high-quality audio using a pre-trained neural vocoder, Parallel WaveGAN. In experiments involving electrocorticography (ECoG) recordings from 13 participants, the synthesized speech achieved Pearson correlation coefficients ranging from 0.85 to 0.95. These results demonstrate the effectiveness of the Transformer decoder in reconstructing accurate spectrograms from ECoG signals, even under data-constrained conditions.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.