Convergent representations and spatiotemporal dynamics of speech and language in brain and deep neural networks
Chen, P.; Xiang, S.; He, L.; Chang, E. F.; Li, Y.
Show abstract
Recent studies have explored the correspondence between single-modality DNN models (speech or text) and specific brain networks for speech and language. The key factors underlying these correlations and their spatiotemporal evolution within the brain language network remain unclear, particularly across different DNN modalities. To address these questions, we analyzed the representation similarity between self-supervised learning (SSL) models for speech (Wav2Vec2) and language (GPT-2), against neural responses to naturalistic speech captured via high-density electrocorticography. Our results indicated high prediction accuracy of both types of SSL models relative to neural activity before and after word onsets. It was the shared components between Wav2Vec2.0 and GPT-2 that explained the majority portion of the SSL-brain similarity. Furthermore, we observed distinct spatiotemporal dynamics: both models showed high encoding accuracy 40 milliseconds before word onset, especially in the mid-superior temporal gyrus (mid-STG), which can be explained by the shared contextual components in the SSL models; the Wav2Vec2.0 also peaked at 200 milliseconds after word onset around the posterior STG, which was mainly attributed to the acoustic-phonetic and static semantic information encoded in the SSL models. These results highlight how contextual and acoustic-phonetic cues encoded in DNNs align with spatiotemporal neural activity patterns, suggesting a significant overlap in how artificial and biological systems process linguistic information.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Parallel hierarchical encoding of linguistic representations in the human auditory cortex and recurrent automatic speech recognition systems 96%
- A Neural Speech Decoding Framework Leveraging Deep Learning and Speech Synthesis 94%
- Accurate and efficient time-domain classification with adaptive spiking recurrent neural networks 93%
Similar papers in this journal
- A Geometric Framework for Understanding Dynamic Information Integration in Context-Dependent Computation 93%
- Unsupervised alignment reveals structural commonalities and differences in neural representations of natural scenes across individuals and brain areas 93%
- The neural representation of visually evoked emotion is high-dimensional, categorical, and distributed across transmodal brain regions 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.