Video representation learning has recently attracted attention in computer\nvision due to its applications for activity and scene forecasting or\nvision-based planning and control. Video prediction models often learn a latent\nrepresentation of video which is encoded from input frames and decoded back\ninto images. Even when conditioned on actions, purely deep learning based\narchitectures typically lack a physically interpretable latent space. In this\nstudy, we use a differentiable physics engine within an action-conditional\nvideo representation network to learn a physical latent representation. We\npropose supervised and self-supervised learning methods to train our network\nand identify physical properties. The latter uses spatial transformers to\ndecode physical states back into images. The simulation scenarios in our\nexperiments comprise pushing, sliding and colliding objects, for which we also\nanalyze the observability of the physical properties. In experiments we\ndemonstrate that our network can learn to encode images and identify physical\nproperties like mass and friction from videos and action sequences in the\nsimulated scenarios. We evaluate the accuracy of our supervised and\nself-supervised methods and compare it with a system identification baseline\nwhich directly learns from state trajectories. We also demonstrate the ability\nof our method to predict future video frames from input images and actions.\n
Paper
References (27)
Scroll for more · 15 remaining