Back

High-level visual cortex representations are organized along visual rather than abstract principles.

Shoham, A.; Broday-Dvir, R.; Malach, R.; Yovel, G.

2025-04-27 neuroscience
10.1101/2024.11.12.623145 bioRxiv
Show abstract

A fundamental question dominating the study of human visual cortex is whether it is organized along visual or semantic information. This question is unresolved, and the controversy has been rekindled by the recent report that surprisingly, revealed that the representations of textual description of images by linguistic artificial networks successfully predict the response of high-level visual cortex to visual images. These findings appear to support a linguistic, abstract, organizing principle of human visual cortex. Here, using iEEG recordings from high level visual cortex in patients, we contributed to this debate, by testing the hypothesis that this linguistic alignment is restricted to textual descriptions of the visual content of the images (visual text) and does not extend to abstract textual descriptions (abstract text). We selected images that depict familiar faces and places, as these images allow for the best dissociation between these two types of text and generated their visual and abstract (e.g., name and biography of a person) textual descriptions. We then predicted the relational structures of the iEEG response to the images using their textual representations based on a large language model and the image representation based on a convolutional neural network. Neural relational-structures in high-level visual cortex were similarly predicted by images and visual-text but not abstract-text representations. Abstract text best predicted responses of the fronto-parietal cortex to the images. These results demonstrate that visual-language alignment in high-level visual cortex is limited to visually grounded language.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.