Divergent specializations for motion-driven representations in higher lateral and dorsal visual areas
Darjani, N.; Bakhtiari, S.; Vaziri-Pashkam, M.; Robert, S.
Show abstract
The human visual system integrates both static and dynamic information to support form and shape perception, yet the computational principles underlying the integration of motion for object recognition remain unclear. Artificial neural networks (ANNs) offer a computational framework for developing and testing hypotheses about these principles: if ANNs trained on motion-related tasks develop representations that align with brain activity and support object categorization, this would suggest that the training objectives and architectural constraints of these networks may capture key aspects of motion processing in biological visual systems in general, and motion processing for object recognition, in particular. Here, we investigated this question using "object kinematograms", stimuli in which object form is conveyed solely through motion cues. We measured neural responses of two higher regions of the lateral and the dorsal visual pathways, respectively, with strong sensitivity to dynamic cues from objects: lateral occipitotemporal cortex (LOTbio), and left supramarginal gyrus (SMGlh), as well as primary visual cortex (V1). We compared brain responses to representations extracted from two neural networks: SlowFast, a dual-pathway architecture trained on action recognition that processes slow- and fast-varying visual information with cross-pathway integration, and DorsalNet, a model of the primate dorsal visual pathway trained on embodied self-motion estimation. Representational similarity analysis revealed distinct representational profiles across brain areas, demonstrating functional specialization in motion-based form processing. LOTbio was best characterized by the slow pathway of the SlowFast model, whereas SMGlh showed strong similarity to both models. Critically, we found that representations aligned with brain activity also better supported behavioral function: the full SlowFast model, incorporating both slow and fast pathways, outperformed other models in few-shot categorization of object kinematograms and showed the highest similarity to human perceptual judgments. These findings demonstrate that with appropriate inductive biases, specifically, dual-pathway architectures for multi-scale motion processing and training objectives focused on dynamic visual tasks, ANNs can develop functionally useful representations of motion-defined forms that exhibit better alignment with the visual regions involved in processing dynamic visual signals.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- How the visual brain can learn to parse images using a multiscale, incremental grouping process 94%
- High-performing neural network models of visual cortex benefit from high latent dimensionality 94%
- Tuned normalization in perceptual decision-making circuits can explain seemingly suboptimal confidence behavior 94%
Similar papers in this journal
- The ventral visual pathway represents animal appearance over animacy, unlike human behavior and deep neural networks 96%
- Predicting the partition of behavioral variability in speed perception with naturalistic stimuli 94%
- The relative coding strength of object identity and nonidentity features in human occipito-temporal cortex and convolutional neural networks 94%
Similar papers in this journal
- Orthogonal Representations of Object Shape and Category in Deep Convolutional Neural Networks and Human Visual Cortex 96%
- Orthogonal neural representations support perceptual judgements of natural stimuli 94%
- Functional characterization of retinal ganglion cells using tailored nonlinear modeling 93%
Similar papers in this journal
- An image-computable model for the stimulus selectivity of gamma oscillations 95%
- Hierarchical temporal prediction captures motion processing from retina to higher visual cortex 95%
- Gain, not concomitant changes in spatial receptive field properties, improves task performance in a neural network attention model 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.