Author Response to Review arMx
We thank the reviewer for helpful, constructive, and positive feedback. We are glad that you find our work "well-motivated", "well-demonstrated", and "outperform baselines on downstream tasks by a large margin". Here, we try to address the questions and concerns:
**Comparison with Baselines on Imitation**
As ASE and CALM do not use a motion-tracking objective for training, we cannot compare motion imitation results directly. We show that they do not have good motion skill coverage in the VR controller tracking task, which is a generalization and arguably harder task than full-body motion imitation. The VR controller tracking task only has partial body observations, but requires the policies to produce the same full-body motion as the full-body imitation task. Between PULSE and imitate \& repurpose, the main difference is how the latent space is trained, where PULSE is through distillation and imitate \& repurpose is through RL. Imitate \& repurpose also does not use a Gaussian prior or learnable prior, though both PULSE and Imitate \& repurpose use the AR1 prior. We report the result of training using RL and no distillation in Table 6 of Appendix C.2. We can see that training without distillation can learn a large amount of motion in AMASS (76\%), but it cannot scale to the performance of PULSE. We hypothesize that random sampling for variational bottleneck and random sampling for RL hampers the learning process. This ablation can be viewed as imitate \& repurpose implemented in our framework and humanoid, and we show that it does not scale as well in the imitation task.
**Comparison between PHC and PHC+**
Both PHC and PHC+ are evaluated on the same set of modified training and testing sequences in Table 1. After removing sequences with issues, there are 11313 and 138 AMASS motion sequences for training and testing, and we re-run all PHC's evaluations on the modified dataset.
**Plot Content Visibility**
We have updated the plot for better visibility. Thanks!
**Answer to Questions**
- As sim2real from physics simulation [1, 2, 3, 4] continues work, we believe that a similar methodology should transfer to real humanoids. Although modifications to the humanoid kinematic tree, reward designs, and domain randomization will be required, the methodology should be applicable. PHC has also demonstrated motion imitation on diverse body shapes, which could also be distilled to a shape/dynamics-aware latent space.
[1] Cheng, Xuxin, Kexin Shi, Ananye Agarwal, and Deepak Pathak. 2023. “Extreme Parkour with Legged Robots.” arXiv [Cs.RO]. arXiv. http://arxiv.org/abs/2309.14341.
[2] Zhuang, Ziwen, Zipeng Fu, Jianren Wang, Christopher Atkeson, Soeren Schwertfeger, Chelsea Finn, and Hang Zhao. 2023. “Robot Parkour Learning.” arXiv [Cs.RO]. arXiv. http://arxiv.org/abs/2309.05665.
[3] Fu, Zipeng, Ashish Kumar, Ananye Agarwal, Haozhi Qi, Jitendra Malik, and Deepak Pathak. 2021. “Coupling Vision and Proprioception for Navigation of Legged Robots.” http://arxiv.org/abs/2112.02094.
[4] Radosavovic, Ilija, Tete Xiao, Bike Zhang, Trevor Darrell, Jitendra Malik, and Koushil Sreenath. 2023. “Learning Humanoid Locomotion with Transformers.” arXiv [Cs.RO], March. https://doi.org/10.48550/ARXIV.2303.03381.