Monkey See, Model Knew: Large Language Models Accurately Predict Visual Brain Responses in Humans and Non-Human Primates
Conwell, C.; MacMahon, E.; Jagadeesh, A.; Vinken, K.; Sharma, S.; Prince, J. S.; Alvarez, G. A.; Konkle, T.; Livingstone, M.; Isik, L.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWRecent progress in multimodal AI and language-aligned visual representation learning has rekindled debates about the role of language in shaping the human visual system. In particular, the emergent ability of language-aligned vision models (e.g. CLIP) - and even pure language models (e.g. BERT) - to predict image-evoked brain activity has led some to suggest that human visual cortex itself may be language-aligned in comparable ways. But what would we make of this claim if the same procedures could model visual activity in a species without language? Here, we conducted controlled comparisons of pure-vision, pure-language, and multimodal vision-language models in their prediction of human (N=4) and rhesus macaque (N=6, 5:IT, 1:V1) ventral visual activity to the same set of 1000 captioned natural images (the NSD1000). The results revealed markedly similar patterns in model predictivity of early and late ventral visual cortex across both species. This suggests that language model predictivity of the human visual system is not necessarily due to the evolution or learning of language perse, but rather to the statistical structure of the visual world that is reflected in natural language.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.