Object-zoomed training of convolutional neural networks inspired by toddler development improves shape bias
Mueller, N.; Snoek, C. G. M.; Groen, I. I. A.; Scholte, H. S.
Show abstract
Convolutional Neural Networks (CNNs) surpass human-level performance on visual object recognition and detection, but their behavior still differs from human behavior in important ways. One prominent example is that CNNs trained on ImageNet exhibit an image texture bias, while humans exhibit a strong bias toward object shape. Although CNN shape bias can be increased in various ways, e.g., using data augmentation or additional training techniques, it remains unclear what causes the strong discrepancy between human and CNN object recognition strategies. Developmental research suggests that one factor driving human shape bias is that during early childhood, toddlers tend to fill their field-of-view with close-up objects. Here, we operationalize this close-up as a zoom-in on objects during CNN training which we show increases shape bias without any additional training or data augmentation. We provide further evidence for the advantage of closeup object vision by systematically manipulating the background-object ratio during CNN training, and demonstrate a strong (inverse) correlation with shape bias. Moreover, zooming-in on objects, thereby more closely emulating child vision, not only increases shape bias but also concurrently aligns classification accuracy and shape bias between humans and CNNs. Finally, we achieve a near human-like shape bias when using a developmentally-inspired background-object ratio for training and shape bias assessment. In sum, from a simple adjustment to common image datasets - zooming-in on objects - human-like shape bias can emerge. These results suggest that taking inspiration from human learning strategies is a promising avenue for building human-aligned, efficient, and more robust vision CNNs.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Constrained sampling from deep generative image models reveals mechanisms of human target detection 95%
- Can deep convolutional neural networks support relational reasoning in the same-different task? 95%
- Deep neural networks trained for estimating albedo and illumination achieve lightness constancy differently than human observers. 92%
Similar papers in this journal
- Leveraging prior concept learning improves ability to generalize from few examples in computational models of human object recognition 95%
- Semantic relatedness emerges in deep convolutional neural networks designed for object recognition 94%
- Hierarchical sparse coding of objects in deep convolutional neural networks 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.