The ability to form complex plans based on raw visual input is a litmus test\nfor current capabilities of artificial intelligence, as it requires a seamless\ncombination of visual processing and abstract algorithmic execution, two\ntraditionally separate areas of computer science. A recent surge of interest in\nthis field brought advances that yield good performance in tasks ranging from\narcade games to continuous control; these methods however do not come without\nsignificant issues, such as limited generalization capabilities and\ndifficulties when dealing with combinatorially hard planning instances. Our\ncontribution is two-fold: (i) we present a method that learns to represent its\nenvironment as a latent graph and leverages state reidentification to reduce\nthe complexity of finding a good policy from exponential to linear (ii) we\nintroduce a set of lightweight environments with an underlying discrete\ncombinatorial structure in which planning is challenging even for humans.\nMoreover, we show that our methods achieves strong empirical generalization to\nvariations in the environment, even across highly disadvantaged regimes, such\nas "one-shot" planning, or in an offline RL paradigm which only provides\nlow-quality trajectories.\n
Paper
References (47)
Scroll for more · 35 remaining