Back

Dynamics of auditory word form encoding in human speech cortex

Zhang, Y.; Leonard, M. K.; Gwilliams, L.; Bhaya-Grossman, I.; Chang, E. F.

2025-05-05 neuroscience
10.1101/2025.05.05.651964 bioRxiv
Show abstract

When we hear continuous speech, we perceive it as a series of discrete words, despite the lack of clear boundaries in the acoustic signal. The superior temporal gyrus (STG) encodes phonetic elements like consonants and vowels, but how it extracts whole words as perceptual units remains unclear. Using high-density cortical recordings, we investigated how the brain represents auditory word forms--integrating acoustic-phonetic, prosodic, and lexical features--while participants listened to spoken narratives. Our results show that STG neural populations exhibit a distinctive reset in activity at word boundaries, marked by a brief, sharp drop in cortical activity. Between these resets, the STG consistently encodes distinct acoustic-phonetic, prosodic, and lexical information, supporting the integration of phonological features into coherent word forms. Notably, this process tracks the relative elapsed time within each word, independent of its absolute duration, providing a flexible temporal scaffolding for encoding variable word lengths. We observed similar word form dynamics in the deeper layers of a self-supervised artificial speech network, suggesting a potential convergence with computational models. Additionally, in a bistable word perception task, STG responses were aligned with participants perceived word boundaries on a trial-by-trial basis, further emphasizing the role of dynamic encoding in word recognition. Together, these findings support a new dynamical model of auditory word forms, highlighting their importance as perceptual units for accessing linguistic meaning.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.