We consider Model-Agnostic Meta-Learning (MAML) methods for Reinforcement\nLearning (RL) problems, where the goal is to find a policy using data from\nseveral tasks represented by Markov Decision Processes (MDPs) that can be\nupdated by one step of stochastic policy gradient for the realized MDP. In\nparticular, using stochastic gradients in MAML update steps is crucial for RL\nproblems since computation of exact gradients requires access to a large number\nof possible trajectories. For this formulation, we propose a variant of the\nMAML method, named Stochastic Gradient Meta-Reinforcement Learning (SG-MRL),\nand study its convergence properties. We derive the iteration and sample\ncomplexity of SG-MRL to find an $\\epsilon$-first-order stationary point, which,\nto the best of our knowledge, provides the first convergence guarantee for\nmodel-agnostic meta-reinforcement learning algorithms. We further show how our\nresults extend to the case where more than one step of stochastic policy\ngradient method is used at test time. Finally, we empirically compare SG-MRL\nand MAML in several deep RL environments.\n
Paper
References (29)
Scroll for more · 17 remaining