Back

Semantic Information Orthogonal to Visual Features Peaks in LateralOccipitotemporal Cortex

Ponnambalam, A. R.; Pottore Venkiteswaran, K.

2026-03-15 neuroscience
10.64898/2026.03.14.711805 bioRxiv
Show abstract

Language model embeddings of scene descriptions predict responses in the human higher visual cortex. However, a fundamental question remains: does this alignment reflect truly visually-independent semantic content, or does it occur because language models better mimic the complex visual features that drive these areas? We used 7T fMRI data from the Natural Scenes Dataset to directly address this by removing the influence of visual feature embeddings from language model embeddings, isolating semantic content that is separate from the visual signal. We then used these visually-independent embeddings to predict brain responses in individual voxels through cross-validated ridge regression. After adjusting for visual signals, we found a clear difference in brain regions: the lateral occipitotemporal cortex, especially in areas selective for body perception, showed significantly more visually-independent semantic variance compared to ventral stream regions. In contrast, the early visual cortex displayed notably negative predictions after adjustment, confirming that our method effectively removed visually-driven signals. This pattern was consistent across all eight subjects, both hemispheres, and six combinations of language models and visual feature architectures. These findings suggest that the lateral stream retains substantially more variance from language models unrelated to various visual feature models than the ventral stream does. This suggests that visually independent semantic coding is organized heterogenously along the occipital cortex. HighlightsO_LIBody-selective lateral occipitotemporal cortex (EBA) contains the strongest visually-independent semantic representations in human visual cortex. C_LIO_LIAfter removing visual feature variance, semantic encoding is significantly greater in lateral stream regions than in canonical ventral stream areas (FFA, PPA, RSC).The lateral-over-ventral dissociation is architecture-invariant, replicating across six combinations of language models (BERT, GPT-2, CLIP-text) and visual feature sets, with GPT-2 > BERT > CLIP-text ordering validating the pipeline. C_LI

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.