A key challenge for an agent learning to interact with the world is to reason\nabout physical properties of objects and to foresee their dynamics under the\neffect of applied forces. In order to scale learning through interaction to\nmany objects and scenes, robots should be able to improve their own performance\nfrom real-world experience without requiring human supervision. To this end, we\npropose a novel approach for modeling the dynamics of a robot's interactions\ndirectly from unlabeled 3D point clouds and images. Unlike previous approaches,\nour method does not require ground-truth data associations provided by a\ntracker or any pre-trained perception network. To learn from unlabeled\nreal-world interaction data, we enforce consistency of estimated 3D clouds,\nactions and 2D images with observed ones. Our joint forward and inverse network\nlearns to segment a scene into salient object parts and predicts their 3D\nmotion under the effect of applied actions. Moreover, our object-centric model\noutputs action-conditioned 3D scene flow, object masks and 2D optical flow as\nemergent properties. Our extensive evaluation both in simulation and with\nreal-world data demonstrates that our formulation leads to effective,\ninterpretable models that can be used for visuomotor control and planning.\nVideos, code and dataset are available at http://hind4sight.cs.uni-freiburg.de\n
Paper
References (41)
Scroll for more · 29 remaining