ECO: Energy-Constrained Optimization with Reinforcement Learning for Humanoid Walking

Achieving stable and energy-efficient locomotion is essential for humanoid robots to operate continuously in real-world applications. Existing model predictive control (MPC) and reinforcement learning (RL) approaches often rely on energy-related metrics embedded within a multi-objective optimization framework, which require extensive hyperparameter tuning and often result in suboptimal policies. To address these challenges, we propose ECO (Energy-Constrained Optimization), a constrained RL framework that separates energy-related metrics from rewards, reformulating them as explicit inequality constraints. This method provides a clear and interpretable physical representation of energy costs, enabling more efficient and intuitive hyperparameter tuning for improved energy efficiency. ECO introduces dedicated constraints for energy consumption and reference motion, enforced by the Lagrangian method, to achieve stable, symmetric, and energy-efficient walking for humanoid robots. We evaluated ECO against MPC, standard RL with reward shaping, and four state-of-the-art constrained RL methods. Experiments, including sim-to-sim and sim-to-real transfers on the kid-sized humanoid robot BRUCE, demonstrate that ECO significantly reduces energy consumption compared to baselines while maintaining robust walking performance. These results highlight a substantial advancement in energy-efficient humanoid locomotion. All experimental demonstrations can be found on the project website: https://sites.google.com/view/eco-humanoid. Note to Practitioners—Traditional MPC and RL approaches often require extensive hyperparameter tuning and frequently result in suboptimal solutions for improving energy efficiency while maintaining stable walking performance. ECO is designed to address these challenges by reformulating energy consumption as explicit inequality constraints, providing a physically interpretable and intuitive approach to optimizing energy efficiency. This framework is particularly well-suited for applications prioritizing energy conservation and operational stability, such as surveillance, disaster response, and long-duration autonomous operations. Additionally, ECO generates emergent behaviors such as lighter steps and reduced body shaking, which are especially advantageous for loco-manipulation tasks by minimizing disruptions to upper-body manipulation caused by locomotion. Comparative experiments empirically offer valuable insights into constraint selection and learning setups, which may also inspire relevant ongoing research in constrained RL.

Paper

Similar papers

© 2026 NYSGPT2525 LLC