Brain-like dynamics in speech representations can emerge through self-supervised learning
Liu, O. D.; Tang, H.; Feldman, N. H.; Goldwater, S.
Show abstract
Speech representations in the human brain do not simply mirror the instantaneous speech signal; rather, they display several properties that are hypothesized to facilitate the integration of speech sounds into words. In particular, neural encodings of speech maintain information that has dissipated from the acoustics, and have also been argued to abstract over variability in how individual speech sounds are produced. Here, we investigate how such characteristics could arise. We introduce a computational framework that uses modern neural network models from speech technology to examine two factors in particular: the learning mechanism and the learning input. We find that self-supervised models trained without lexical or semantic feedback developed temporal dynamics similar to brain representations, regardless of whether they were trained on speech or non-speech audio. In contrast, only models trained on speech learn to abstract over variability due to word position and phonetic context. Overall, our results suggest that domain-general learning mechanisms can lead to several important properties of speech representations, but in some cases require domain-specific input in order to do so. Significance StatementUnderstanding speech often feels effortless, but in fact mapping speech sounds into words involves complex computation. Experimental neuroscience has identified key properties in brain signals that may support this computation, but why and how these properties arise is still unclear. We examined these properties in computational models and found that they occurred in models that werent given lexical or semantic feedback, but were trained to predict the acoustics of the speech signal. This suggests such properties can develop from domain-general learning combined with domain-specific input. Moreover, some properties even arose in models that were trained on non-speech audio. Overall, our work illustrates how computational modeling can help reveal the conditions under which neural properties emerge.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Deep neural network models reveal interplay of peripheral coding and stimulus statistics in pitch perception 95%
- Convergent neural signatures of speech prediction error are a biological marker for spoken word recognition 94%
- The impact of musical expertise on disentangled and contextual neural encoding of music revealed by generative music models 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.