Foundation models like Vision-Language Models (VLMs) excel at commonsense vision and language tasks such as visual question answering. However, they cannot yet directly solve complex, long-horizon robot manipulation problems requiring precise continuous reasoning. Task and Motion Planning (TAMP) systems can handle long-horizon reasoning through discrete-continuous hybrid search over parameterized skills, but rely on detailed environment models and cannot interpret novel human objectives, such as arbitrary natural language goals. We propose integrating VLMs into TAMP systems by having them generate continuous (and discrete) language-parameterized constraints that enable open-world reasoning. Specifically, we use VLMs to generate continuous action constraints in the form of code, which augment traditional TAMP manipulation constraints. We also generate discrete plan constraints that parameterize actions and partially order them with respect to the continuous constraints. Experiments show that our approach, OWL-TAMP, outperforms ablations relying solely on TAMP or VLMs, as well as several VLM planning baselines from the literature, across several long-horizon manipulation tasks specified directly in natural language. We additionally demonstrate that OWL-TAMP can be deployed with an off-the-shelf TAMP system to solve challenging manipulation tasks on real-world hardware.
Paper
References (95)
Scroll for more · 38 remaining