Back

Large Language Models Reveal the Neural Tracking of Linguistic Context in Attended and Unattended Multi-Talker Speech

Puffay, C.; Mischler, G.; Choudhari, V.; Vanthornhout, J.; Bickel, S.; Mehta, A. D.; Schevon, C. A.; McKhann, G. M.; Van hamme, H.; Francart, T.; Mesgarani, N.

2025-04-24 neuroscience
10.1101/2025.04.24.648897 bioRxiv
Show abstract

Large language models (LLMs) capture long-range contextual structure in natural language and have recently been shown to align with the human brains contextualized linguistic encoding. This makes them a promising computational probe for studying how context-dependent linguistic information is represented during natural speech perception. Speech perception often occurs in multi-talker environments, where attention must dynamically select among competing streams, yet how contextual information from attended and unattended speech is neurally encoded remains underexplored. Here, we investigate how auditory attention modulates neural tracking of context-dependent linguistic representations using electrocorticography (ECoG) and stereoelectroencephalography (sEEG) recordings from three epilepsy patients engaged in a two-conversation "cocktail party" paradigm. To model neural responses to attended and unattended speech streams, we used contextual word embeddings generated by large language models. We find that LLM-derived features reliably predict brain activity for the attended stream and that contextual information from the unattended stream also contributes to neural prediction. Importantly, these contributions extend beyond low-level acoustic features and shallow syntactic information, and depend on the surrounding linguistic context. Moreover, neural tracking of the unattended stream reflects shorter-range contextual integration than that of the attended stream. Together, these findings indicate that neural responses to speech reflect context-dependent linguistic representations from multiple concurrent speech streams, with attention modulating the depth and timescale of contextual integration. Our results highlight the utility of LLMs for probing higher-level linguistic representations in complex, naturalistic listening environments.

Published in Imaging Neuroscience (predicted rank #1) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.