Despite a series of recent successes in reinforcement learning (RL), many RL\nalgorithms remain sensitive to hyperparameters. As such, there has recently\nbeen interest in the field of AutoRL, which seeks to automate design decisions\nto create more general algorithms. Recent work suggests that population based\napproaches may be effective AutoRL algorithms, by learning hyperparameter\nschedules on the fly. In particular, the PB2 algorithm is able to achieve\nstrong performance in RL tasks by formulating online hyperparameter\noptimization as time varying GP-bandit problem, while also providing\ntheoretical guarantees. However, PB2 is only designed to work for continuous\nhyperparameters, which severely limits its utility in practice. In this paper\nwe introduce a new (provably) efficient hierarchical approach for optimizing\nboth continuous and categorical variables, using a new time-varying bandit\nalgorithm specifically designed for the population based training regime. We\nevaluate our approach on the challenging Procgen benchmark, where we show that\nexplicitly modelling dependence between data augmentation and other\nhyperparameters improves generalization.\n
Paper
References (96)
Scroll for more · 38 remaining