Predictive vision-language integration in the human visual cortex
Li, S.; Jin, Z.; Zhang, R.-Y.; Gu, S.; Li, Y.
Show abstract
Integrating linguistic and visual information is a core function of human cognition, yet how information from these two modalities interacts in the brain remains largely unknown. Competing frameworks, including the hub-and-spoke model and Bayesian theories such as predictive coding, offer conflicting accounts of how the brain achieves multimodal integration. To address this question, we collected a large-scale fMRI dataset and leveraged state-of-the-art AI systems to construct encoding models that probe how the human brain matches and integrates linguistic and visual information. We found that prior information from one modality can modulate neural responses in another, even in the early visual cortex (EVC). Integration neural response in EVC is governed by prediction errors consistent with predictive coding theory. Enhanced and suppressed neural responses to semantically matched cross-modal stimuli were found in distinct EVC populations, with suppression population carrying denser, behaviorally relevant semantic information. Both populations support semantic integration with distinct temporal dynamics and representational structures. These findings provide representational- and computational-level insights into how the brain integrates information across modalities, revealing unified principles of information processing that link biological and artificial intelligence.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Spatial processing of limbs reveals the center-periphery bias in high level visual cortex follows a nonlinear topography 98%
- Dissociable contributions of the medial parietal cortex to recognition memory 97%
- Neural evidence for boundary updating as the source of the repulsive bias in classification 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.