Back

A two-dimensional space of linguistic representations shared across individuals

Tuckute, G.; Lee, E. J.; Ou, Y.; Fedorenko, E.; Kay, K.

2025-05-23 neuroscience
10.1101/2025.05.21.655330 bioRxiv
Show abstract

Our ability to extract meaning from linguistic inputs and package ideas into word sequences is supported by a network of left-hemisphere frontal and temporal brain areas. Despite extensive research, previous attempts to discover differences among these language areas have not revealed clear dissociations or spatial organization. All areas respond similarly during controlled linguistic experiments as well as during naturalistic language comprehension. To search for finer-grained organizational principles of language processing, we applied data-driven decomposition methods to ultra-high-field (7T) fMRI responses from eight participants listening to 200 linguistically diverse sentences. Using a cross-validation procedure that identifies shared structure across individuals, we find that two components successfully generalize across participants, together accounting for about 32% of the explainable variance in brain responses to sentences. The first component corresponds to processing difficulty, and the second--to meaning abstractness; we formally support this interpretation through targeted behavioral experiments and information-theoretic measures. Furthermore, we find that the two components are systematically organized within frontal and temporal language areas, with the meaning-abstractness component more prominent in the temporal regions. These findings reveal an interpretable, low-dimensional, spatially structured representational basis for language processing, and advance our understanding of linguistic representations at a detailed, fine-scale organizational level.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.