A Probabilistic Interpretation of Self-Paced Learning with Applications to Reinforcement Learning

Across machine learning, the use of curricula has shown strong empirical\npotential to improve learning from data by avoiding local optima of training\nobjectives. For reinforcement learning (RL), curricula are especially\ninteresting, as the underlying optimization has a strong tendency to get stuck\nin local optima due to the exploration-exploitation trade-off. Recently, a\nnumber of approaches for an automatic generation of curricula for RL have been\nshown to increase performance while requiring less expert knowledge compared to\nmanually designed curricula. However, these approaches are seldomly\ninvestigated from a theoretical perspective, preventing a deeper understanding\nof their mechanics. In this paper, we present an approach for automated\ncurriculum generation in RL with a clear theoretical underpinning. More\nprecisely, we formalize the well-known self-paced learning paradigm as inducing\na distribution over training tasks, which trades off between task complexity\nand the objective to match a desired task distribution. Experiments show that\ntraining on this induced distribution helps to avoid poor local optima across\nRL algorithms in different tasks with uninformative rewards and challenging\nexploration requirements.\n

Paper

References (85)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC