Back

Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural language

Hosseini, E. A.; Fedorenko, E.

2023-11-06 neuroscience
10.1101/2023.11.05.564832 bioRxiv
Show abstract

Predicting upcoming events is critical to our ability to effectively interact with our environment and conspecifics. In natural language processing, transformer models, which are trained on next-word prediction, appear to construct a general-purpose representation of language that can support diverse downstream tasks. However, we still lack an understanding of how a predictive objective shapes such representations. Inspired by recent work in vision neuroscience Henaff et al. (2019), here we test a hypothesis about predictive representations of autoregressive transformer models. In particular, we test whether the neural trajectory of a sequence of words in a sentence becomes progressively more straight as it passes through the layers of the network. The key insight behind this hypothesis is that straighter trajectories should facilitate prediction via linear extrapolation. We quantify straightness using a 1-dimensional curvature metric, and present four findings in support of the trajectory straightening hypothesis: i) In trained models, the curvature progressively decreases from the first to the middle layers of the network. ii) Models that perform better on the next-word prediction objective, including larger models and models trained on larger datasets, exhibit greater decreases in curvature, suggesting that this improved ability to straighten sentence neural trajectories may be the underlying driver of better language modeling performance. iii) Given the same linguistic context, the sequences that are generated by the model have lower curvature than the ground truth (the actual continuations observed in a language corpus), suggesting that the model favors straighter trajectories for making predictions. iv) A consistent relationship holds between the average curvature and the average surprisal of sentences in the middle layers of models, such that sentences with straighter neural trajectories also have lower surprisal. Importantly, untrained models dont exhibit these behaviors. In tandem, these results support the trajectory straightening hypothesis and provide a possible mechanism for how the geometry of the internal representations of autoregressive models supports next word prediction.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 1%
18.3%
2
Neurobiology of Language
29 papers in training set
Top 0.1%
18.3%
3
Nature Human Behaviour
95 papers in training set
Top 0.1%
9.7%
4
Scientific Reports
3612 papers in training set
Top 17%
5.5%
50% of probability mass above
5
eLife
5828 papers in training set
Top 22%
5.4%
6
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 9%
5.4%
7
Nature Communications
5641 papers in training set
Top 33%
4.0%
8
eneuro
439 papers in training set
Top 3%
2.7%
9
Journal of Neurophysiology
302 papers in training set
Top 1%
2.6%
10
Psychological Review
19 papers in training set
Top 0.1%
2.4%
11
Neural Computation
39 papers in training set
Top 0.4%
2.1%
12
Neural Networks
35 papers in training set
Top 0.3%
2.1%
13
Cognition
47 papers in training set
Top 0.4%
1.7%
14
Neuron
337 papers in training set
Top 4%
1.7%
15
iScience
1154 papers in training set
Top 19%
1.7%
16
The Journal of Neuroscience
1025 papers in training set
Top 7%
1.7%
17
Nature Machine Intelligence
70 papers in training set
Top 2%
1.0%
18
Journal of Cognitive Neuroscience
135 papers in training set
Top 2%
1.0%
19
Cerebral Cortex
396 papers in training set
Top 5%
0.8%
20
Journal of Vision
110 papers in training set
Top 0.7%
0.8%
21
NeuroImage
903 papers in training set
Top 6%
0.8%
22
PLOS ONE
5266 papers in training set
Top 62%
0.8%
23
Communications Psychology
22 papers in training set
Top 0.5%
0.6%