Computational Counterfactuals Reveal Non-Additive Audiovisual Semantics in Natural Movie Responses
Li, M.
Show abstract
Natural audiovisual perception may not be fully captured by decomposing movies into auditory and visual streams. I introduce a computational-counterfactual framework that keeps movie viewing intact while varying only AI-derived descriptions of the same clips. Using 7 Tesla movie fMRI imaging data from 176 participants, I tested whether cortical responses were better predicted by native audiovisual semantics than by a dimension-matched additive reconstruction from audio-only and video-only descriptions. The native model outperformed the matched additive baseline under content-aware purged cross-validation, with strongest gains in auditory, visual, and dorsal attention systems. Representational-similarity, feature-replacement, and content-gating analyses showed that the advantage reflected feature- and network-specific routing linked to coherent audiovisual semantic emergence rather than raw auditory-visual discrepancy. The effect survived stronger temporal purging and repeat-content exclusion, suggesting that intact movie viewing evokes cortical structure aligned with native audiovisual meaning beyond additive unimodal semantics.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Two common and distinct forms of variation in human functional brain networks 96%
- Structure and influence in an interconnected world: neurocomputational mechanism of real-time distributed learning on social networks 96%
- Interplay between persistent activity and activity-silent dynamics in prefrontal cortex during working memory 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.