The Surprising Effectiveness of Linear Models for Visual Foresight in Object Pile Manipulation
In this paper, we tackle the problem of pushing piles of small objects into a\ndesired target set using visual feedback. Unlike conventional single-object\nmanipulation pipelines, which estimate the state of the system parametrized by\npose, the underlying physical state of this system is difficult to observe from\nimages. Thus, we take the approach of reasoning directly in the space of\nimages, and acquire the dynamics of visual measurements in order to synthesize\na visual-feedback policy. We present a simple controller using an image-space\nLyapunov function, and evaluate the closed-loop performance using three\ndifferent class of models for image prediction: deep-learning-based models for\nimage-to-image translation, an object-centric model obtained from treating each\npixel as a particle, and a switched-linear system where an action-dependent\nlinear map is used. Through results in simulation and experiment, we show that\nfor this task, a linear model works surprisingly well -- achieving better\nprediction error, downstream task performance, and generalization to new\nenvironments than the deep models we trained on the same amount of data. We\nbelieve these results provide an interesting example in the spectrum of models\nthat are most useful for vision-based feedback in manipulation, considering\nboth the quality of visual prediction, as well as compatibility with rigorous\nmethods for control design and analysis. Project site:\nhttps://sites.google.com/view/linear-visual-foresight/home\n