Scale matters: Large language models with billions (rather than millions) of parameters better match neural representations of natural language
Hong, Z.; Wang, K.; Zada, Z.; Gazula, H.; Turner, D.; Aubrey, B.; Niekerken, L.; Doyle, W.; Devore, S.; Dugan, P.; Friedman, D.; Devinsky, O.; Flinker, A.; Hasson, U.; Nastase, S.; Goldstein, A.
Show abstract
Recent research has used large language models (LLMs) to study the neural basis of naturalistic language processing in the human brain. LLMs have rapidly grown in complexity, leading to improved language processing capabilities. However, neuroscience researchers havent kept up with the quick progress in LLM development. Here, we utilized several families of transformer-based LLMs to investigate the relationship between model size and their ability to capture linguistic information in the human brain. Crucially, a subset of LLMs were trained on a fixed training set, enabling us to dissociate model size from architecture and training set size. We used electrocorticography (ECoG) to measure neural activity in epilepsy patients while they listened to a 30-minute naturalistic audio story. We fit electrode-wise encoding models using contextual embeddings extracted from each hidden layer of the LLMs to predict word-level neural signals. In line with prior work, we found that larger LLMs better capture the structure of natural language and better predict neural activity. We also found a log-linear relationship where the encoding performance peaks in relatively earlier layers as model size increases. We also observed variations in the best-performing layer across different brain regions, corresponding to an organized language processing hierarchy.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Artificial neural network language models predict human brain responses to language even after a developmentally realistic amount of training 96%
- Lexical semantic content, not syntactic structure, is the main contributor to ANN-brain similarity of fMRI responses in the language network 95%
- Beyond Letters: Optimal Transport as a Model for Sub-Letter Orthographic Processing 95%
Similar papers in this journal
- Representational similarity learning reveals a graded multi-dimensional semantic space in the human anterior temporal cortex 96%
- Encoding neural representations of time-continuous stimulus-response transformations in the human brain with advanced deep neural networks 95%
- Semantic composition in experimental and naturalistic paradigms 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.