Learning to See Like a Child: Why Viewpoint Diversity is Fundamental for Human-Aligned Object Recognition
Luo, Y.; Müller, N.; Scholte, H. S.
Show abstract
Deep convolutional neural networks match human accuracy on standard object recognition tasks but fail to recognize familiar objects from novel view-points. Humans, however, develop viewpoint-invariant recognition at an early age through diverse visual experience. This gap in visual experience may explain why models diverge from humans in object recognition. Holding dataset size constant, we show that greater viewpoint diversity substantially improves generalization to novel views. Using a synthetic 3D dataset with systematically controlled viewpoints, we reveal a core trade-off: restricted-view training yields rapid learning and near-ceiling in-distribution accuracy but collapses on held-out viewpoints, whereas viewpoint-diverse training learns more gradually yet generalizes robustly. Increasing viewpoint diversity disrupts texture regularities while preserving global shape, driving networks to prioritize shape over texture - the same strategy that underlies human object recognition. Partitioned Grad-CAM analyses further show that viewpoint-diverse models maintain object-centered attention. These findings parallel developmental accounts of multi-view learning and identify viewpoint diversity as an important factor for robust, human-aligned vision.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Constrained sampling from deep generative image models reveals mechanisms of human target detection 97%
- Deep neural networks trained for estimating albedo and illumination achieve lightness constancy differently than human observers. 94%
- Can deep convolutional neural networks support relational reasoning in the same-different task? 94%
Similar papers in this journal
- Leveraging prior concept learning improves ability to generalize from few examples in computational models of human object recognition 95%
- Semantic relatedness emerges in deep convolutional neural networks designed for object recognition 93%
- The face module emerged in a deep convolutional neural network selectively deprived of face experience 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.