Speaker Identity is Robustly Encoded in Spatial Patterns of Intracranial EEG for Attention Decoding
Dindar, S. S.; Jiang, X.; Choudhari, V.; Bickel, S.; Mehta, A.; Schevon, C.; McKhann, G. M.; Friedman, D.; Flinker, A.; Mesgarani, N.
Show abstract
The human auditory cortex robustly tracks attended speech, yet it remains unclear if speaker identity is encoded in spatial patterns of neural activity independent of temporal dynamics. Here, we demonstrate that the identity of an attended speaker is reliably reflected in distinct, time-invariant spatial activation maps in human intracranial EEG (iEEG). Leveraging these "neural fingerprints", we developed a novel framework for Auditory Attention Decoding (AAD) that shifts from traditional temporal envelope tracking to spatial speaker identification. By decoupling the decoding of "who" is speaking from "when" they are speaking, our modular system achieves state-of-the-art speech extraction, particularly in short time windows (<2 seconds) where temporal models typically fail. Furthermore, we observed a reciprocal shift in neural activity during attentional switches, confirming that these spatial codes dynamically track listener intent. These findings establish that speaker identity is a robust, spatially distributed feature in the auditory cortex, offering a high-speed, complementary mechanism for neuro-steered hearing technologies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- The Neural Response at the Fundamental Frequency of Speech is Modulated by Word-level Acoustic and Linguistic Information 94%
- Cortical tracking of voice pitch in the presence of multiple speakers depends on selective attention 93%
- Decoding age-related changes in the spatiotemporal neural processing of speech using machine learning 93%
Similar papers in this journal
- Neural speech tracking benefit of lip movements predicts behavioral deterioration when the speaker's mouth is occluded 94%
- Sensory and perceptual decisional processes underlying the perception of reverberant auditory environments 94%
- Dynamic time-locking mechanism in the cortical representation of spoken words 93%
Similar papers in this journal
- Alignment massive of auditory individual artificial networks with fMRI brain data leads to generalizable improvements in brain encoding and downstream tasks 95%
- Harmonizing and aligning M/EEG datasets with covariance-based techniques to enhance predictive regression modeling 94%
- Dynamic Network Analysis of Electrophysiological Task Data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.