Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning

Despite the close connection between exploration and sample efficiency, most\nstate of the art reinforcement learning algorithms include no considerations\nfor exploration beyond maximizing the entropy of the policy. In this work we\naddress this seeming missed opportunity. We observe that the most common\nformulation of directed exploration in deep RL, known as bonus-based\nexploration (BBE), suffers from bias and slow coverage in the few-sample\nregime. This causes BBE to be actively detrimental to policy learning in many\ncontrol tasks. We show that by decoupling the task policy from the exploration\npolicy, directed exploration can be highly effective for sample-efficient\ncontinuous control. Our method, Decoupled Exploration and Exploitation Policies\n(DEEP), can be combined with any off-policy RL algorithm without modification.\nWhen used in conjunction with soft actor-critic, DEEP incurs no performance\npenalty in densely-rewarding environments. On sparse environments, DEEP gives a\nseveral-fold improvement in data efficiency due to better exploration.\n

Paper

References (51)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC