One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RL

While reinforcement learning algorithms can learn effective policies for\ncomplex tasks, these policies are often brittle to even minor task variations,\nespecially when variations are not explicitly provided during training. One\nnatural approach to this problem is to train agents with manually specified\nvariation in the training task or environment. However, this may be infeasible\nin practical situations, either because making perturbations is not possible,\nor because it is unclear how to choose suitable perturbation strategies without\nsacrificing performance. The key insight of this work is that learning diverse\nbehaviors for accomplishing a task can directly lead to behavior that\ngeneralizes to varying environments, without needing to perform explicit\nperturbations during training. By identifying multiple solutions for the task\nin a single environment during training, our approach can generalize to new\nsituations by abandoning solutions that are no longer effective and adopting\nthose that are. We theoretically characterize a robustness set of environments\nthat arises from our algorithm and empirically find that our diversity-driven\napproach can extrapolate to various changes in the environment and task.\n

Paper

References (47)

Scroll for more · 35 remaining

Similar papers

© 2026 NYSGPT2525 LLC