Visually-Grounded Planning without Vision: Language Models Infer Detailed Plans from High-level Instructions

The recently proposed ALFRED challenge task aims for a virtual robotic agent\nto complete complex multi-step everyday tasks in a virtual home environment\nfrom high-level natural language directives, such as "put a hot piece of bread\non a plate". Currently, the best-performing models are able to complete less\nthan 5% of these tasks successfully. In this work we focus on modeling the\ntranslation problem of converting natural language directives into detailed\nmulti-step sequences of actions that accomplish those goals in the virtual\nenvironment. We empirically demonstrate that it is possible to generate gold\nmulti-step plans from language directives alone without any visual input in 26%\nof unseen cases. When a small amount of visual information is incorporated,\nnamely the starting location in the virtual environment, our best-performing\nGPT-2 model successfully generates gold command sequences in 58% of cases. Our\nresults suggest that contextualized language models may provide strong visual\nsemantic planning modules for grounded virtual agents.\n

Paper

References (19)

Scroll for more · 7 remaining

Similar papers

© 2026 NYSGPT2525 LLC