Correspondence between the layered structure of deep language models and temporal structure of natural language processing in the human brain
Goldstein, A.; Ham, E.; Nastase, S. A.; Zada, Z.; Dabush, A.; Bobbi Aubrey, B.; Schain, M.; Gazula, H.; Feder, A.; Doyle, W.; Devore, S.; Dugan, P.; Friedman, D.; Brenner, M.; Hassidim, A.; Devinsky, O.; Flinker, A.; Levy, O.; Hasson, U.
Show abstract
Deep language models (DLMs) provide a novel computational paradigm for how the brain processes natural language. Unlike symbolic, rule-based models described in psycholinguistics, DLMs encode words and their context as continuous numerical vectors. These "embeddings" are constructed by a sequence of computations organized in "layers" to ultimately capture surprisingly sophisticated representations of linguistic structures. How does this layered hierarchy map onto the human brain during natural language comprehension? In this study, we used electrocorticography (ECoG) to record neural activity in language areas along the superior temporal gyrus and inferior frontal gyrus while human participants listened to a 30-minute spoken narrative. We supplied this same narrative to a high-performing DLM (GPT2-XL) and extracted the contextual embeddings for each word in the story across all 48 layers of the model. We next trained a set of linear encoding models to predict the temporally-evolving neural activity from the embeddings at each layer. We found a striking correspondence between the layer-by-layer sequence of embeddings from GPT2-XL and the temporal sequence of neural activity in language areas. In addition, we found evidence for the gradual accumulation of recurrent information along the linguistic processing hierarchy. However, we also noticed additional neural processes in the brain, but not in DLMs, during the processing of surprising (unpredictable) words. These findings point to a connection between human language processing and DLMs where the layer-by-layer accumulation of contextual information in DLM embeddings matches the temporal dynamics of neural activity in high-order language areas.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The impact of musical expertise on disentangled and contextual neural encoding of music revealed by generative music models 96%
- Shared functional specialization in transformer-based language models and the human brain 96%
- Incremental Accumulation of Linguistic Context in Artificial and Biological Neural Networks 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.