Off-policy deep reinforcement learning algorithms commonly compensate for\noverestimation bias during temporal-difference learning by utilizing\npessimistic estimates of the expected target returns. In this work, we propose\nGeneralized Pessimism Learning (GPL), a strategy employing a novel learnable\npenalty to enact such pessimism. In particular, we propose to learn this\npenalty alongside the critic with dual TD-learning, a new procedure to estimate\nand minimize the magnitude of the target returns bias with trivial\ncomputational cost. GPL enables us to accurately counteract overestimation bias\nthroughout training without incurring the downsides of overly pessimistic\ntargets. By integrating GPL with popular off-policy algorithms, we achieve\nstate-of-the-art results in both competitive proprioceptive and pixel-based\nbenchmarks.\n
Paper
References (67)
Scroll for more · 38 remaining