A generalizable speech neuroprosthesis
Fogg, Z. M.; Card, N. S.; Wairagkar, M.; Srinivasan, A.; Singer-Clark, T.; Hou, X.; Okorokova, E.; Peracha, H.; Iacobacci, C.; Brailow, T.; Jude, J. J.; Levi-Aharoni, H.; Le, T.; Mifsud, D.; Deevi, P.; Nason-Tomaszewski, S.; Pritchard, A. L.; Zhang, Y.; Richards, B.; Bechefsky, P.; Hochberg, L. R.; Williams, Z.; Shahlaie, K.; Au Yong, N.; Rubin, D.; Pandarinath, C.; Brandman, D. M.; Stavisky, S. D.
Show abstract
Intracortical brain-computer interfaces (BCIs) can restore communication to people with vocal tract paralysis by decoding cortical activity during attempted speech into text. State-of-the-art systems pairing neural-to-phoneme decoders with phoneme-to-word language models have achieved word error rates (WERs) as low as 1%, but only after collecting thousands of sentences of training data. Shortening the data collection process would facilitate scaling this new technology by reducing the time from device implant to high-accuracy communication. Here we introduce a transformer-based decoder model trained jointly across six intracortical speech BCI participants. For every participant -- regardless of sex, disease etiology, or attempted speaking strategy -- a multi-user model decoded speech more accurately (over 50% lower relative WER on average) than models trained on individual users data. Notably, the multi-user model could be finetuned on fewer than 200 sentences from a held-out user to achieve a WER below 7%. These results reveal how to pool intracortical data across people to yield more accurate, generalizable, and rapidly-deployable decoding models.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Ultra-fast Prediction of Somatic Structural Variations by Reduced Read Mapping via Pan-Genome k-mer Sets 91%
- Identifying perturbations that boost T-cell infiltration into tumours via counterfactual learning of their spatial proteomic profiles 90%
- Biomimetic multi-channel microstimulation of somatosensory cortex conveys high resolution force feedback for bionic hands 90%
Similar papers in this journal
- Parallel hierarchical encoding of linguistic representations in the human auditory cortex and recurrent automatic speech recognition systems 94%
- Accurate and efficient time-domain classification with adaptive spiking recurrent neural networks 94%
- A Neural Speech Decoding Framework Leveraging Deep Learning and Speech Synthesis 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.