Back

Vision-language encoding models reveal an image-computable food-quality dimension in human occipitotemporal cortex

Marrazzo, G.; Pimpini, L.; Roefs, A.

2026-08-28 neuroscience
10.64898/2026.08.25.747006 bioRxiv
Show abstract

Perceived calorie content contributes to neural representational structure in human ventral visual cortex, yet it remains unclear whether this reflects an abstract nutritional signal or whether perceived calorie is largely recoverable from the visual-semantic structure of the food image itself. In 25 female participants who passively viewed 96 food images during functional MRI, we decomposed perceived calorie ratings into a component predicted from CLIP (Contrastive Language-Image Pretraining) image embeddings, a vision-language model that captures high-level visual-semantic image structure, and a residual component not captured by this CLIP-based prediction. We then tested their respective contributions to neural prediction using cross-validated banded ridge encoding models. The CLIP-predictable component organized foods along a processedness and naturalness dimension, separating raw single-ingredient foods from prepared and energy-dense foods. Adding this component to a visual-semantic baseline improved neural prediction progressively along the ventral visual hierarchy, with the strongest relative contribution in higher-level ventral temporal cortex. These findings indicate that calorie-related encoding in ventral visual cortex is carried mainly by a shared food-quality axis indexing processedness, naturalness, and perceived healthiness, a substantial part of which is recoverable from image-computable visual-semantic structure, rather than providing evidence for an isolated abstract representation of caloric magnitude. Because these results derive from a reanalysis of 25 female participants viewing a fixed set of 96 images, generalization to broader populations and larger, more varied stimulus sets remains to be established.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.