Systematic image perturbations reveal persistent gaps between human and machine vision
Kato, M.; He, B. J.
Show abstract
Deep neural networks (DNNs) are promising computational models for understanding visual object recognition. Yet, whether DNNs use similar visual cues for object recognition as humans do remains unknown. We created an image set that systematically untangles global shape, internal parts, and texture information, and compared human recognition behavior against >200 DNNs spanning diverse architectures, training diets, and training objectives. No DNNs replicated humans cue-reliance profile, including those with recurrence or specialized training. Fine-tuned text-image contrastive-trained models, regardless of architecture, were most human-like overall, but lost their human-alignment when the global shape was disrupted. Strikingly, all DNNs substantially underperformed humans when the global shape cue alone was critical to object recognition. Furthermore, alignment with ventral stream neural recordings in an existing database did not predict alignment to human behavior, and model performance does not always predict its human-alignment. Together, these findings reveal systematic and persistent differences between human and machine vision.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Neural state space alignment for magnitude generalization in humans and recurrent networks 94%
- Inhibitory and excitatory populations in parietal cortex are equally selective for decision outcome in both novices and experts 94%
- When the ventral visual stream is not enough: A deep learning account of medial temporal lobe involvement in perception 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.