While visual imitation learning offers one of the most effective ways of\nlearning from visual demonstrations, generalizing from them requires either\nhundreds of diverse demonstrations, task specific priors, or large,\nhard-to-train parametric models. One reason such complexities arise is because\nstandard visual imitation frameworks try to solve two coupled problems at once:\nlearning a succinct but good representation from the diverse visual data, while\nsimultaneously learning to associate the demonstrated actions with such\nrepresentations. Such joint learning causes an interdependence between these\ntwo problems, which often results in needing large amounts of demonstrations\nfor learning. To address this challenge, we instead propose to decouple\nrepresentation learning from behavior learning for visual imitation. First, we\nlearn a visual representation encoder from offline data using standard\nsupervised and self-supervised learning methods. Once the representations are\ntrained, we use non-parametric Locally Weighted Regression to predict the\nactions. We experimentally show that this simple decoupling improves the\nperformance of visual imitation models on both offline demonstration datasets\nand real-robot door opening compared to prior work in visual imitation. All of\nour generated data, code, and robot videos are publicly available at\nhttps://jyopari.github.io/VINN/.\n
Paper
References (53)
Scroll for more · 38 remaining