Joint Representation of Color and Shape in Convolutional Neural Networks: A Stimulus-rich Network Perspective
Taylor, J.; Xu, Y.
Show abstract
To interact with real-world objects, any effective visual system must jointly code the unique features defining each object. Despite decades of neuroscience research, we still lack a firm grasp on how the primate brain binds visual features. Here we apply a novel network-based stimulus-rich representational similarity approach to study color and shape binding in five convolutional neural networks (CNNs) with varying architecture, depth, and presence/absence of recurrent processing. All CNNs showed near-orthogonal color and shape processing in early layers, but increasingly interactive feature coding in higher layers, with this effect being much stronger for networks trained for object classification than untrained networks. These results characterize for the first time how multiple visual features are coded together in CNNs. The approach developed here can be easily implemented to characterize whether a similar coding scheme may serve as a viable solution to the binding problem in the primate brain.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Constrained sampling from deep generative image models reveals mechanisms of human target detection 96%
- How the forest interacts with the trees: Multiscale shape integration explains global and local processing 94%
- Biased orientation representations can be explained by experience with non-uniform training set statistics 94%
Similar papers in this journal
Similar papers in this journal
- Gain, not concomitant changes in spatial receptive field properties, improves task performance in a neural network attention model 97%
- Emergent Color Categorization in a Neural Network trained for Object Recognition 96%
- Increasing stimulus similarity drives nonmonotonic representational change in hippocampus 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.