Revisiting Peng's Q($λ$) for Modern Reinforcement Learning

Off-policy multi-step reinforcement learning algorithms consist of\nconservative and non-conservative algorithms: the former actively cut traces,\nwhereas the latter do not. Recently, Munos et al. (2016) proved the convergence\nof conservative algorithms to an optimal Q-function. In contrast,\nnon-conservative algorithms are thought to be unsafe and have a limited or no\ntheoretical guarantee. Nonetheless, recent studies have shown that\nnon-conservative algorithms empirically outperform conservative ones. Motivated\nby the empirical results and the lack of theory, we carry out theoretical\nanalyses of Peng's Q($\\lambda$), a representative example of non-conservative\nalgorithms. We prove that it also converges to an optimal policy provided that\nthe behavior policy slowly tracks a greedy policy in a way similar to\nconservative policy iteration. Such a result has been conjectured to be true\nbut has not been proven. We also experiment with Peng's Q($\\lambda$) in complex\ncontinuous control tasks, confirming that Peng's Q($\\lambda$) often outperforms\nconservative algorithms despite its simplicity. These results indicate that\nPeng's Q($\\lambda$), which was thought to be unsafe, is a theoretically-sound\nand practically effective algorithm.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC