Representing Multiple Visual Objects in the Human Brain and Convolutional Neural Networks
Mocz, V.; Jeong, S.; Chun, M.; Xu, Y.
Show abstract
Objects in the real world often appear with other objects. To recover the identity of an object whether or not other objects are encoded concurrently, in primate object-processing regions, neural responses to an object pair have been shown to be well approximated by the average responses to each constituent object shown alone, indicating the whole is equal to the average of its parts. This is present at the single unit level in the slope of response amplitudes of macaque IT neurons to paired and single objects, and at the population level in response patterns of fMRI voxels in human ventral object processing regions (e.g., LO). Here we show that averaging exists in both single fMRI voxels and voxel population responses in human LO, with better averaging in single voxels leading to better averaging in fMRI response patterns, demonstrating a close correspondence of averaging at the fMRI unit and population levels. To understand if a similar averaging mechanism exists in convolutional neural networks (CNNs) pretrained for object classification, we examined five CNNs with varying architecture, depth and the presence/absence of recurrent processing. We observed averaging at the CNN unit level but rarely at the population level, with CNN unit response distribution in most cases did not resemble human LO or macaque IT responses. The whole is thus not equal to the average of its parts in CNNs, potentially rendering the individual objects in a pair less accessible in CNNs during visual processing than they are in the human brain.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting identity-preserving object transformations across the human ventral visual stream 97%
- The relative coding strength of object identity and nonidentity features in human occipito-temporal cortex and convolutional neural networks 97%
- The spatiotemporal neural dynamics of object recognition for natural images and line drawings 97%
Similar papers in this journal
- The locus of recognition memory signals in human cortex depends on the complexity of the memory representations 96%
- Scene-selective brain regions respond to embedded objects of a scene 96%
- Individual differences in spatial working memory strategies differentially reflected in the engagement of control and default brain networks 95%
Similar papers in this journal
- Understanding transformation tolerant visual object representations in the human brain and convolutional neural networks 98%
- Temporal contiguity training does not affect size-tolerant representations in object-selective cortex 97%
- The contribution of object size, manipulability, and stability on neural responses to inanimate objects 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.