Decoding of the speech envelope from EEG using the VLAAI deep neural network.
Accou, B.; Vanthornhout, J.; Van hamme, H.; Francart, T.
Show abstract
To investigate the processing of speech in the brain, commonly simple linear models are used to establish a relationship between brain signals and speech features. However, these linear models are ill-equipped to model a highly-dynamic, complex non-linear system like the brain, and they often require a substantial amount of subject-specific training data. This work introduces a novel speech decoder architecture: the Very Large Augmented Auditory Inference (VLAAI) network. The VLAAI network outperformed state-of-the-art subject-independent models (median Pearson correlation of 0.19, p < 0.001), yielding an increase over the well-established linear model by 52%. Using ablation techniques we identified the relative importance of each part of the VLAAI network and found that the non-linear components and output context module influenced model performance the most (10% relative performance increase). Subsequently, the VLAAI network was evaluated on a holdout dataset of 26 subjects and publicly available unseen dataset to test generalization for unseen subjects and stimuli. No significant difference was found between the holdout subjects and the default test set, and only a small difference between the default test set and the public dataset was found. Compared to the baseline models, the VLAAI network still significantly outperformed all baseline models on the public dataset. We evaluated the effect of training set size by training the VLAAI network on data from 1 up to 80 subjects and evaluated on 26 holdout subjects, revealing a logarithmic relationship between the number of subjects in the training set and the performance on unseen subjects. Finally, the subject-independent VLAAI network was fine-tuned for 26 holdout subjects to obtain subject-specific VLAAI models. With 5 minutes of data or more, a significant performance improvement was found, up to 34% (from 0.18 to 0.25 median Pearson correlation) with regards to the subject-independent VLAAI network.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- The Neural Response at the Fundamental Frequency of Speech is Modulated by Word-level Acoustic and Linguistic Information 96%
- Speech-driven Facial Animations Improve Speech-in-Noise Comprehension of Humans 95%
- Decoding continuous variables from event-related potential (ERP) data with linear support vector regression (SVR) using the Decision Decoding Toolbox (DDTBOX) 95%
Similar papers in this journal
- Online speech synthesis using a chronically implanted brain-computer interface in an individual with ALS 96%
- Feasibility of decoding covert speech in ECoG with aTransformer trained on overt speech 95%
- Comparison of Two-Talker Attention Decoding from EEG with Nonlinear Neural Networks and Linear Methods 95%
Similar papers in this journal
- Alignment massive of auditory individual artificial networks with fMRI brain data leads to generalizable improvements in brain encoding and downstream tasks 96%
- Harmonizing and aligning M/EEG datasets with covariance-based techniques to enhance predictive regression modeling 95%
- Surfing beta burst waveforms to improve motor imagery-based BCI 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.