Toward Non-Invasive Voice Restoration: A Deep Learning Approach Using Real-Time MRI
Saleh, M. W.
Show abstract
Despite recent advances in brain-computer interfaces (BCIs) for speech restoration, existing systems remain invasive, costly, and inaccessible to individuals with congenital mutism or neurodegenerative disease. We present a proof-of-concept pipeline that synthesizes personalized speech directly from real-time magnetic resonance imaging (rtMRI) of the vocal tract, without requiring acoustic input. Segmented rtMRI frames are mapped to articulatory class representations using a Pix2Pix conditional GAN, which are then transformed into synthetic audio waveforms by a convolutional neural network modeling the articulatory-to-acoustic relationship. The outputs are rendered into audible form and evaluated with speaker-similarity metrics derived from Resemblyzer embeddings. While preliminary, our results suggest that even silent articulatory motion encodes sufficient information to approximate a speakers vocal characteristics, offering a non-invasive direction for future speech restoration in individuals who have lost or never developed voice.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FONDUE: Robust resolution-invariant denoising of MR Images using Nested UNets 94%
- EPISeg: Automated segmentation of the spinal cord on echo planar images using open-access multi-center data 93%
- Alignment massive of auditory individual artificial networks with fMRI brain data leads to generalizable improvements in brain encoding and downstream tasks 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.