Dear reviewer RQdG,
Thank you for your thoughtful comments and feedback! Here, we respond to specific questions and concerns the reviewer raises.
---------------------------------------
**Question/Comment 1**: It would be nice to see LIV performance on the ALOHA dataset.
**Response 1**: Thanks for this suggestion. We have now added an updated Figure 3 to include comparison between GVL (Zero-Shot) vs. LIV on our ALOHA-250 dataset. As shown, LIV’s qualitative behavior is similar to the one on OXE – the VOCs are close to a uniform distribution. In conclusion, GVL still produces task progress estimates that are higher quality than LIV on ALOHA-250, though the gap has become smaller compared to the results on OXE since the ALOHA dataset is sufficiently challenging for GVL too.
---------------------------------------
**Question/Comment 2**: The use of VOC on the OXE evaluation is somewhat circular.
**Response 2**: Thanks for this suggestion. To break up the circularity, we need to manually delegate a few datasets to be of “high-quality” to use VOC as a metric. To this end, we have selected RT-1 [1], Bridge [2, 3], DOBBE [4], and BC-Z [5] datasets because these datasets have successfully been used in prior works for large-scale imitation learning. The result is illustrated in Figure 10, Appendix F of our updated manuscript. As seen, on this subset of high-quality and diverse OXE datasets, GVL consistently generates highly positive VOCs; in contrast, LIV generates VOCs close to a uniform distribution, indicating that it struggles to accurately predict language-conditioned task progress on unseen robot videos. These results validate our experimental protocol of using VOCs as a metric for universal value functions. In addition, we have also improved the writing in Section 4.1 to make the distinction clearer.
[1] Brohan, Anthony, et al. "Rt-1: Robotics transformer for real-world control at scale." arXiv preprint arXiv:2212.06817 (2022).
[2] Ebert, Frederik, et al. "Bridge data: Boosting generalization of robotic skills with cross-domain datasets." arXiv preprint arXiv:2109.13396 (2021).
[3] Walke, Homer Rich, et al. "Bridgedata v2: A dataset for robot learning at scale." Conference on Robot Learning. PMLR, 2023.
[4] Shafiullah, Nur Muhammad Mahi, et al. "On bringing robots home." arXiv preprint arXiv:2311.16098 (2023).
[5] Jang, Eric, et al. "Bc-z: Zero-shot task generalization with robotic imitation learning." Conference on Robot Learning. PMLR, 2022.
---------------------------------------
**Question/Comment 3**: more analysis regarding GVL’s ability to accurately estimate values in videos where repetition is present.
**Response 3**: Thanks for this suggestion! We have included two videos in the updated supplementary material in which repetition or failure is present. As shown, in both cases, when the robot gripper pulls back from the object of task interest, the GVL task progress estimates decrease, and when the gripper recovers and makes progress again, GVL estimates increase. These results demonstrate GVL’s ability to accurately estimate values in videos when repetition is present.
---------------------------------------
**Question/Comment 4**: ALOHA-13 should be clarified.
**Response 4**: Thank you for this suggestion. We have added a sentence in Section 4.2 clarifying. In short, ALOHA-13 is a subset of ALOHA-250 for which we have collected additional demonstration data. For example, instead of having only 2-3 demonstrations per task in ALOHA-250, we have 500+ Demos per task in ALOHA-13. This allows us to obtain statistically significant results on a task level for the in-context value learning experiments.
We thank the reviewer again for their time and effort helping us improve our paper! Please let us know if we can provide additional clarifications to improve our score.