GradientDICE: Rethinking Generalized Offline Estimation of Stationary Values

We present GradientDICE for estimating the density ratio between the state\ndistribution of the target policy and the sampling distribution in off-policy\nreinforcement learning. GradientDICE fixes several problems of GenDICE (Zhang\net al., 2020), the state-of-the-art for estimating such density ratios. Namely,\nthe optimization problem in GenDICE is not a convex-concave saddle-point\nproblem once nonlinearity in optimization variable parameterization is\nintroduced to ensure positivity, so any primal-dual algorithm is not guaranteed\nto converge or find the desired solution. However, such nonlinearity is\nessential to ensure the consistency of GenDICE even with a tabular\nrepresentation. This is a fundamental contradiction, resulting from GenDICE's\noriginal formulation of the optimization problem. In GradientDICE, we optimize\na different objective from GenDICE by using the Perron-Frobenius theorem and\neliminating GenDICE's use of divergence. Consequently, nonlinearity in\nparameterization is not necessary for GradientDICE, which is provably\nconvergent under linear function approximation.\n

Paper

References (44)

Scroll for more · 32 remaining

Similar papers

© 2026 NYSGPT2525 LLC