Visual Attention Through Uncertainty Minimization in Recurrent Generative Models
Standvoss, K.; Quax, S. C.; van Gerven, M. A. J.
Show abstract
Allocating visual attention through saccadic eye movements is a key ability of intelligent agents. Attention is both influenced through bottom-up stimulus properties as well as top-down task demands. The interaction of these two attention mechanisms is not yet fully understood. A parsimonious reconciliation posits that both processes serve the minimization of predictive uncertainty. We propose a recurrent generative neural network model that predicts a visual scene based on foveated glimpses. The model shifts its attention in order to minimize the uncertainty in its predictions. We show that the proposed model produces naturalistic eye movements focusing on informative stimulus regions. Introducing additional tasks modulates the saccade patterns towards task-relevant stimulus regions. The models saccade characteristics correspond well with previous experimental data in humans, providing evidence that uncertainty minimization could be a fundamental mechanisms for the allocation of visual attention.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Constrained sampling from deep generative image models reveals mechanisms of human target detection 97%
- Target Identification Under High Levels of Amplitude, Size, Orientation and Background Uncertainty 96%
- Can deep convolutional neural networks support relational reasoning in the same-different task? 96%
Similar papers in this journal
- The influence of stereopsis on visual saliency in a proto-object based model of selective attention 96%
- Crowding Reveals Fundamental Differences in Local vs. Global Processing in Humans and Machines 95%
- Visual shape discrimination in goldfish, modelled with the neural circuitry of optic tectum and torus longitudinalis. 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.