A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
Many reinforcement learning algorithms rely on value estimation, however, the\nmost widely used algorithms -- namely temporal difference algorithms -- can\ndiverge under both off-policy sampling and nonlinear function approximation.\nMany algorithms have been developed for off-policy value estimation based on\nthe linear mean squared projected Bellman error (MSPBE) and are sound under\nlinear function approximation. Extending these methods to the nonlinear case\nhas been largely unsuccessful. Recently, several methods have been introduced\nthat approximate a different objective -- the mean-squared Bellman error (MSBE)\n-- which naturally facilitate nonlinear approximation. In this work, we build\non these insights and introduce a new generalized MSPBE that extends the linear\nMSPBE to the nonlinear setting. We show how this generalized objective unifies\nprevious work and obtain new bounds for the value error of the solutions of the\ngeneralized objective. We derive an easy-to-use, but sound, algorithm to\nminimize the generalized objective, and show that it is more stable across\nruns, is less sensitive to hyperparameters, and performs favorably across four\ncontrol domains with neural network function approximation.\n
Paper
References (62)
Scroll for more · 38 remaining