Context and Attention Shape Electrophysiological Correlates of Speech-to-Language Transformation
Anderson, A. J.; Davis, C.; Lalor, E. C.
Show abstract
To transform speech into words, the human brain must accommodate variability across utterances in intonation, speech rate, volume, accents and so on. A promising approach to explaining this process has been to model electroencephalogram (EEG) recordings of brain responses to speech. Contemporary models typically invoke speech categories (e.g. phonemes) as an intermediary representational stage between sounds and words. However, such categorical models are typically hand-crafted and therefore incomplete because they cannot speak to the neural computations that putatively underpin categorization. By providing end-to-end accounts of speech-to-language transformation, new deep-learning systems could enable more complete brain models. We here model EEG recordings of audiobook comprehension with the deep-learning system Whisper. We find that (1) Whisper provides an accurate, self-contained EEG model of speech-to-language transformation; (2) EEG modeling is more accurate when including prior speech context, which pure categorical models do not support; (3) EEG signatures of speech-to-language transformation depend on listener-attention.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Semantic composition in experimental and naturalistic paradigms 95%
- Are you talking to me? How the choice of speech register impacts listeners' hierarchical encoding of speech. 95%
- Alignment massive of auditory individual artificial networks with fMRI brain data leads to generalizable improvements in brain encoding and downstream tasks 95%
Similar papers in this journal
- Expectations boost the reconstruction of auditory features from electrophysiological responses to noisy speech 96%
- Attentional Modulation of Hierarchical Speech Representations in a Multitalker Environment 96%
- Lateralised cerebral processing of abstract linguistic structure in clear and degraded speech 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.