Recurrent Attention Models with Object-centric Capsule Representation for Multi-object Recognition

The visual system processes a scene using a sequence of selective glimpses,\neach driven by spatial and object-based attention. These glimpses reflect what\nis relevant to the ongoing task and are selected through recurrent processing\nand recognition of the objects in the scene. In contrast, most models treat\nattention selection and recognition as separate stages in a feedforward\nprocess. Here we show that using capsule networks to create an object-centric\nhidden representation in an encoder-decoder model with iterative glimpse\nattention yields effective integration of attention and recognition. We\nevaluate our model on three multi-object recognition tasks; highly overlapping\ndigits, digits among distracting clutter and house numbers, and show that it\nlearns to effectively move its glimpse window, recognize and reconstruct the\nobjects, all with only the classification as supervision. Our work takes a step\ntoward a general architecture for how to integrate recurrent object-centric\nrepresentation into the planning of attentional glimpses.\n

Paper

References (68)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC