On the Convergence of Approximate and Regularized Policy Iteration Schemes

Entropy regularized algorithms such as Soft Q-learning and Soft Actor-Critic,\nrecently showed state-of-the-art performance on a number of challenging\nreinforcement learning (RL) tasks. The regularized formulation modifies the\nstandard RL objective and thus generally converges to a policy different from\nthe optimal greedy policy of the original RL problem. Practically, it is\nimportant to control the sub-optimality of the regularized optimal policy. In\nthis paper, we establish sufficient conditions for convergence of a large class\nof regularized dynamic programming algorithms, unified under regularized\nmodified policy iteration (MPI) and conservative value iteration (VI) schemes.\nWe provide explicit convergence rates to the optimality depending on the\ndecrease rate of the regularization parameter. Our experiments show that the\nempirical error closely follows the established theoretical convergence rates.\nIn addition to optimality, we demonstrate two desirable behaviours of the\nregularized algorithms even in the absence of approximations: robustness to\nstochasticity of environment and safety of trajectories induced by the policy\niterates.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC