Summary
The paper presents Hierarchical Planning with Foundation Models (HiP), a framework that addresses long-horizon decision-making in novel environments, using multiple modalities of data and reasoning at various levels of hierarchy. The key elements of this framework involve the use of a large language model for constructing symbolic plans, a video diffusion model to generate observation trajectory plans, and a large pre-trained robot ego-centric action model to map these plans to a robot's control space. The authors propose an iterative refinement approach for feedback incorporation, promoting a consensus among different models, without the need for large model finetuning. Experimental results on two long-horizon tabletop manipulation tasks were presented, demonstrating promising results for the proposed strategy. The authors also mention the potential for including other modalities, like touch and sound, in the future.
Strengths
The technical content of the paper appears to be sound, and the proposed framework has shown promising results in the experimental results provided. The methodology to combine language models, video diffusion models, and ego-centric robot control models into one system is comprehensive.
The paper is well-structured and clearly explains the approach and the reasoning behind the decisions made. The paper's flow from problem statement to proposed solution, and finally to experimental evaluation, is logical and easy to follow.
Weaknesses
1. The baselines compared are all relatively weak. It's more appropriate to compare with foundation robotics models, such as SayCan, Gato[1], palm-e[2]. The authors argue that compared to saycan, they can generalise to new skills. However, they didn't demonstrate the generalization to new skills either. (They evaluated on unseen tasks but not new skills.)
2. While the iterative refinement approach is interesting, it could be further scrutinized to understand its limitations better, particularly concerning computational efficiency and robustness.
3. As this is mostly an experimental, please provide code to ensure reproducibility.
4. Formatting issue in line 9, 45.
[1] Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez,
Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian
Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Nando de Freitas. A generalist agent. In Transactions on Machine Learning
Research (TMLR), November 10, 2022.
[2] Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson,
Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke,
Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. Palm-e: An embodied multimodal
language model. In arXiv preprint arXiv:2303.03378, 2023.