A brain-inspired object-based attention network for multi-object recognition and visual reasoning
Adeli, H.; Ahn, S.; Zelinsky, G.
Show abstract
The visual system uses sequences of selective glimpses to objects to support goal-directed behavior, but how is this attention control learned? Here we present an encoder-decoder model inspired by the interacting bottom-up and top-down visual pathways making up the recognitionattention system in the brain. At every iteration, a new glimpse is taken from the image and is processed through the "what" encoder, a hierarchy of feedforward, recurrent, and capsule layers, to obtain an object-centric (object-file) representation. This representation feeds to the "where" decoder, where the evolving recurrent representation provides top-down attentional modulation to plan subsequent glimpses and impact routing in the encoder. We demonstrate how the attention mechanism significantly improves the accuracy of classifying highly overlapping digits. In a visual reasoning task requiring comparison of two objects, our model achieves near-perfect accuracy and significantly outperforms larger models in generalizing to unseen stimuli. Our work demonstrates the benefits of object-based attention mechanisms taking sequential glimpses of objects.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Small hand-designed convolutional neural networks outperform transfer learning in automated cell shape detection in confluent tissues 95%
- Deep learning models for COVID-19 chest x-ray classification: Preventing shortcut learning using feature disentanglement 95%
- Deep Learning Classification of Lipid Droplets in Quantitative Phase Images 94%
Similar papers in this journal
Similar papers in this journal
- Enhanced cell segmentation with limited annotated data using generative adversarial networks 94%
- Enhanced Cell Tracking Using A GAN-based Super-Resolution Video-to-Video Time-Lapse Microscopy Generative Model 94%
- Coherently Remapping Toroidal Cells But Not Grid Cells are Responsible for Path Integration in Virtual Agents 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.