A recurrent vision transformer shows signatures of primate visual attention

We present a Recurrent Vision Transformer (Recurrent ViT) that integrates a capacity-limited spatial memory module with self-attention to emulate primate-like visual attention. Trained via reinforcement learning on a spatially cued orientation-change detection task, our model exhibits hallmark behavioral signatures of primate attention—including improved detection accuracy and faster reaction times for cued stimuli that scale with cue validity. Analysis of its self-attention maps reveals rich temporal dynamics: spatial biases induced by cues are maintained during blank intervals and reactivated prior to anticipated stimulus changes, mirroring the top–down modulation observed in primate studies. Moreover, targeted manipulations of internal attention weights yield performance changes analogous to those produced by microstimulation in attentional control regions such as the frontal eye fields and superior colliculus. These findings demonstrate that embedding recurrent, memory-driven mechanisms within transformer architectures may provide a computational framework for linking artificial and biological attention

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC