Back

Brain-like dynamics in speech representations can emerge through self-supervised learning

Liu, O. D.; Tang, H.; Feldman, N. H.; Goldwater, S.

2026-01-17 neuroscience
10.64898/2026.01.16.700011 bioRxiv
Show abstract

Speech representations in the human brain do not simply mirror the instantaneous speech signal; rather, they display several properties that are hypothesized to facilitate the integration of speech sounds into words. In particular, neural encodings of speech maintain information that has dissipated from the acoustics, and have also been argued to abstract over variability in how individual speech sounds are produced. Here, we investigate how such characteristics could arise. We introduce a computational framework that uses modern neural network models from speech technology to examine two factors in particular: the learning mechanism and the learning input. We find that self-supervised models trained without lexical or semantic feedback developed temporal dynamics similar to brain representations, regardless of whether they were trained on speech or non-speech audio. In contrast, only models trained on speech learn to abstract over variability due to word position and phonetic context. Overall, our results suggest that domain-general learning mechanisms can lead to several important properties of speech representations, but in some cases require domain-specific input in order to do so. Significance StatementUnderstanding speech often feels effortless, but in fact mapping speech sounds into words involves complex computation. Experimental neuroscience has identified key properties in brain signals that may support this computation, but why and how these properties arise is still unclear. We examined these properties in computational models and found that they occurred in models that werent given lexical or semantic feedback, but were trained to predict the acoustics of the speech signal. This suggests such properties can develop from domain-general learning combined with domain-specific input. Moreover, some properties even arose in models that were trained on non-speech audio. Overall, our work illustrates how computational modeling can help reveal the conditions under which neural properties emerge.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.