Correlation-based binocular disparity computations induce representational bottlenecks at the population level
Wundari, B. G.; Fujita, I.; Ban, H.
Show abstract
Binocular stereopsis depends on comparing the images seen by the two eyes. Although correlation-based models explain responses of individual binocular neurons in primary visual cortex (V1), it remains elusive whether such computations can support depth perception at the population level. Using psychophysics, fMRI, and deep neural networks, we investigated human stereopsis with dynamic anticorrelated stimuli that were dominated by mismatched binocular information. Humans reliably perceived reversed depth as predicted by correlation-based computations, yet population representations consistent with this percept emerged in mid-dorsal V3A, not in V1. Similarly, correlation-based neural networks failed to reproduce human depth judgments. Superposition theory from AI interpretability analysis reveals that correlation-based networks represented features with strong entanglement in shared dimensions, leading to destructive interference. Conversely, architectures integrating non-correlation mechanisms exhibited less entangled representations, aligning closely with human behavior. These findings suggest that correlation mechanisms induce representational bottlenecks at the population level, requiring the joint contribution of correlation and non-correlation processing channels to support robust stereopsis. Significance StatementThe brain must infer depth from binocular inputs that are inherently ambiguous. Although correlation-based models explain disparity tuning of individual neurons in primary visual cortex (V1), whether these local mechanisms support perceptual inference at the population level remains unclear. Using psychophysics, fMRI, and neural network modeling, we show that population representations consistent with perceived depth under ambiguity emerge in mid-dorsal area V3A, not V1. Analyses of how neural networks encode multiple features reveal that correlation-based computations represent features into overlapping activity patterns, producing signal cross-talk that degrades depth estimates. In contrast, models incorporating non-correlation computations maintain more distinct population codes and better match human perception. These findings indicate that correlation-based processing alone limits population representations and that robust stereopsis requires complementary mechanisms.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Probing the Link Between Vision and Language in Material Perception UsingPsychophysics and Unsupervised Learning 95%
- Mechanisms of human dynamic object recognition revealed by sequential deep neural networks 95%
- Diverse and flexible behavioral strategies arise in recurrent neural networks trained on multisensory decision making 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.