The Human Brain as a Dynamic Mixture of Expert Models in Video Understanding
Sartzetaki, C.; Zonneveld, A. W.; Oyarzo, P.; Gifford, A. T.; Cichy, R. M.; Mettes, P.; Groen, I. I.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWThe human brain is the most efficient and versatile system for processing dynamic visual input. By comparing representations from deep video models to brain activity, we can gain insights into mechanistic solutions for effective video processing, important to better understand the brain and to build better models. Current works in model-brain alignment primarily focus on fMRI measurements, leaving open questions about fine-grained dynamic processing. Here, we introduce the first large-scale model benchmarking on alignment to dynamic electroencephalography (EEG) recordings of short natural videos. We analyze 100+ models across the axes of temporal integration, classification task, architecture, and pretraining, using our proposed Cross-Temporal Representational Similarity Analysis (CT-RSA) which matches the best time-unfolded model features to dynamically evolving brain responses, distilling 107 alignment scores. Our findings reveal novel insights on how continuous visual input is integrated in the brain, beyond the standard temporal processing hierarchy from low to high-level representations. After initial alignment to hierarchical static object processing, responses in posterior electrodes best align to mid-level temporally-integrative action features, showing high temporal correspondence to feature timings. In contrast, responses in frontal electrodes best align with high-level static action representations and show no temporal correspondence to the video. Additionally, temporally-integrating state-space models show superior alignment to intermediate posterior activity, in which self-supervised pretraining is also beneficial. We draw a metaphor to a dynamic mixture of expert models for the changing neural preference in tasks and temporal integration reflected in the alignment to different model types across time. We posit that a single best-aligned model would need such training and architecture as to allow combining and dynamically switching between these capacities.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Encoding neural representations of time-continuous stimulus-response transformations in the human brain with advanced deep neural networks 96%
- Alignment massive of auditory individual artificial networks with fMRI brain data leads to generalizable improvements in brain encoding and downstream tasks 96%
- The Individualized Neural Tuning Model: Precise and generalizable cartography of functional architecture in individual brains 96%
Similar papers in this journal
Similar papers in this journal
- Evidence for transient, uncoupled power and functional connectivity dynamics 96%
- DeepComBat: A Statistically Motivated, Hyperparameter-Robust, Deep Learning Approach to Harmonization of Neuroimaging Data 95%
- Prediction of individual melodic contour processing in sensory association cortices from resting state functional connectivity 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.