Delayed Geometric Discounts: An Alternative Criterion for Reinforcement Learning

of autonomous capable of complex theoretical background to optimal In on geometric discounts to evaluate this optimality. processes where future returns are not less valuable. Depending the sample-inefficiency are decayed) mecha-nisms (to deal with sparse, deceptive or adversarial rewards). In this paper, we tackle these issues by generalizing the discounted problem formulation with a family of delayed objective functions. We investigate the underlying RL problem to derive: 1) the optimal stationary solution and 2) an approximation of the optimal non-stationary control. The devised algorithms solved hard exploration problems on tabular environment and improved sample-efficiency on classic simulated robotics benchmarks.

Paper

Similar papers

© 2026 NYSGPT2525 LLC