Online reinforcement learning (RL) methods are often data-inefficient or unreliable, making them difficult to train on real robotic hardware, especially quadruped robots. So learning robotic tasks from pre-collected data is a promising direction. Agile and stable legged locomotion remains an open issue in its general form. Analogous to the rapid progress of supervised learning in recent years, the combination of offline reinforcement learning (ORL) and realistic datasets has the potential to make breakthroughs in this challenging field. To facilitate the ORL research for real-world applications, we benchmark ten ORL algorithms in the realistic quadrupedal locomotion dataset. The dataset is collected by the classical model predictive control (MPC) method, rather than the online RL method commonly utilized by previous ORL benchmarks. Extensive experimental results show that the best-performing ORL algorithms can achieve competitive performance compared with the online RL, and even surpass it in some tasks. However, there is still a gap between the learning-based methods and classical MPC, especially in terms of stability and task response accuracy. Our benchmark can provide a fertile ground for future application-oriented ORL research.
Paper
References (38)
Scroll for more · 26 remaining