Robust Decoding of Speech Acoustics from EEG: Going Beyond the Amplitude Envelope
MacIntyre, A. D.; Gaultier, C.; Goehring, T.
Show abstract
ObjectiveDuring speech perception, properties of the acoustic stimulus can be reconstructed from the listeners brain using methods such as electroencephalography (EEG). Most studies employ the amplitude envelope as a target for decoding; however, speech acoustics can be characterised on multiple dimensions, including as spectral descriptors. The current study assesses how robustly an extended acoustic feature set can be decoded from EEG under varying levels of intelligibility and acoustic clarity. ApproachAnalysis was conducted using EEG from 38 young adults who heard intelligible and non-intelligible speech that was either unprocessed or spectrally degraded using vocoding. We extracted a set of acoustic features which, alongside the envelope, characterised instantaneous properties of the speech spectrum (e.g., spectral slope) or spectral change over time (e.g., spectral flux). We establish the robustness of feature decoding by employing multiple model architectures and, in the case of linear decoders, by standardising decoding accuracy (Pearsons r) using randomly permuted surrogate data. Main resultsLinear models yielded the highest r relative to non-linear models. However, the separate decoder architectures produced a similar pattern of results across features and experimental conditions. After converting r values to Z-scores scaled by random data, we observed substantive differences in the noise floor between features. Decoding accuracy significantly varies by spectral degradation and speech intelligibility for some features, but such differences are reduced in the most robustly decoded features. This suggests acoustic feature reconstruction is primarily driven by generalised auditory processing. SignificanceOur results demonstrate that linear decoders perform comparably to non-linear decoders in capturing the EEG response to speech acoustic properties beyond the amplitude envelope, with the reconstructive accuracy of some features also associated with understanding and spectral clarity. This sheds light on how sound properties are differentially represented by the brain and shows potential for clinical applications moving forward.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Online speech synthesis using a chronically implanted brain-computer interface in an individual with ALS 95%
- Ocular dynamics reveal articulatory processing at single-phoneme level during silent reading 95%
- Audio-visual combination of syllables involves time-sensitive dynamics following from fusion failure 95%
Similar papers in this journal
- Hearing and cognitive decline in aging differentially impact neural tracking of context-supported versus random speech across linguistic timescales 95%
- Predictors for Estimating Subcortical EEG Responses to Continuous Speech 95%
- Disentangling listening effort and memory load beyond behavioural evidence 94%
Similar papers in this journal
- Attention Decoding at the Cocktail Party: Preserved in Hearing Aid Users, Reduced in Cochlear Implant Users 96%
- Decoding of selective attention to continuous speech from the human auditory brainstem response 96%
- The integration of continuous audio and visual speech in a cocktail-party environment depends on attention 96%
Similar papers in this journal
- A comparison of EEG encoding models using audiovisual stimuli and their unimodal counterparts 96%
- Convolutional neural networks can identify brain interactions involved in decoding spatial auditory attention 96%
- Time-resolved dynamic computational modeling of human EEG recordings reveals gradients of generative mechanisms for the MMN response 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.