Three-dimensional shape cues affect human and artificial recognition systems differently
Cutler, M.; Baumel, L.; Tocco, J.; Friebel, W.; Thiruvathukal, G. K.; Baker, N.
Show abstract
Humans and neural networks use shape and texture information differently. While shape is the primary cue in human object recognition, neural networks are more biased towards texture cues. Many tests of shape vs. texture bias have focused on shape recognition from an objects external contour. However, shape information is also conveyed through internal contours, shading, and attached shadows, especially when an object is viewed from noncanonical perspectives. Using models from ShapeNet, we created datasets of 120,000 texture-substituted images of objects from many viewpoints with and without shading and attached shadows. We tested humans and several neural networks ability to classify these objects by both their shape and their texture. Humans were much better at classifying texture-substituted objects by their shape than any network, although these differences were greater when shape was defined only by the external contour than when 3D cues were included. Our findings suggest that networks texture bias is reduced when 3D cues are included in images. We next tested whether the inclusion of 3D cues benefitted humans and neural networks more for images of objects viewed from canonical or noncanonical perspectives. Consistent with earlier research, we found that 3D cues primarily benefitted humans for noncanonical images. For neural networks, the greatest performance gains were for canonical images. These findings suggest fundamental differences in how humans and networks use shading and attached shadows for object recognition. We argue that humans use these cues to infer objects 3D structures while neural networks use them as another surface-level cue like texture.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Constrained sampling from deep generative image models reveals mechanisms of human target detection 96%
- Can deep convolutional neural networks support relational reasoning in the same-different task? 95%
- Local cues enable classification of image patches as surfaces, object boundaries, or illumination changes 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.