A recurrent vision transformer shows signatures of primate visual attention
Morgan, J. M.; Albanna, B.; Herman, J. P.
Show abstract
We present a Recurrent Vision Transformer (Recurrent ViT) that integrates a capacity-limited spatial memory module with self-attention to emulate primate-like visual attention. Trained via reinforcement learning on a spatially cued orientation-change detection task, our model exhibits hallmark behavioral signatures of primate attention--including improved detection accuracy and faster reaction times for cued stimuli that scale with cue validity. Analysis of its self-attention maps reveals rich temporal dynamics: spatial biases induced by cues are maintained during blank intervals and reactivated prior to anticipated stimulus changes, mirroring the top-down modulation observed in primate studies. Moreover, targeted manipulations of internal attention weights yield performance changes analogous to those produced by microstimulation in attentional control regions such as the frontal eye fields and superior colliculus. These findings demonstrate that embedding recurrent, memory-driven mechanisms within transformer architectures may provide a computational framework for linking artificial and biological attention
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Arousal state affects perceptual decision-making by modulating hierarchical sensory processing in a large-scale visual system model 97%
- Mechanisms of human dynamic object recognition revealed by sequential deep neural networks 97%
- Novelty is not Surprise: Human exploratory and adaptive behavior in sequential decision-making 97%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.