We present a slot-wise, object-based transition model that decomposes a scene\ninto objects, aligns them (with respect to a slot-wise object memory) to\nmaintain a consistent order across time, and predicts how those objects evolve\nover successive frames. The model is trained end-to-end without supervision\nusing losses at the level of the object-structured representation rather than\npixels. Thanks to its alignment module, the model deals properly with two\nissues that are not handled satisfactorily by other transition models, namely\nobject persistence and object identity. We show that the combination of an\nobject-level loss and correct object alignment over time enables the model to\noutperform a state-of-the-art baseline, and allows it to deal well with object\nocclusion and re-appearance in partially observable environments.\n
Paper
References (40)
Scroll for more · 28 remaining