Multiplicative Controller Fusion: Leveraging Algorithmic Priors for Sample-efficient Reinforcement Learning and Safe Sim-To-Real Transfer

Learning-based approaches often outperform hand-coded algorithmic solutions\nfor many problems in robotics. However, learning long-horizon tasks on real\nrobot hardware can be intractable, and transferring a learned policy from\nsimulation to reality is still extremely challenging. We present a novel\napproach to model-free reinforcement learning that can leverage existing\nsub-optimal solutions as an algorithmic prior during training and deployment.\nDuring training, our gated fusion approach enables the prior to guide the\ninitial stages of exploration, increasing sample-efficiency and enabling\nlearning from sparse long-horizon reward signals. Importantly, the policy can\nlearn to improve beyond the performance of the sub-optimal prior since the\nprior's influence is annealed gradually. During deployment, the policy's\nuncertainty provides a reliable strategy for transferring a simulation-trained\npolicy to the real world by falling back to the prior controller in uncertain\nstates. We show the efficacy of our Multiplicative Controller Fusion approach\non the task of robot navigation and demonstrate safe transfer from simulation\nto the real world without any fine-tuning. The code for this project is made\npublicly available at https://sites.google.com/view/mcf-nav/home\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC