Offline Meta Reinforcement Learning with Weighted Policy Constraints and Proximal Context Collection
Offline meta-reinforcement learning (OMRL) encounters two key challenges: effectively learning the meta-policy from offline datasets and correctly inferring unseen tasks. Existing methods often address the first challenge by imposing policy constraints, but are limited by the suboptimal actions in offline datasets. For the second challenge, most focus on meta-training without enhancing task inference during meta-testing. To address these issues, we propose a novel method called weighted policy conStraints and proximal contExt coLlECtion sTrategy for OMRL (SELECT). During meta-training, we integrate policy constraints with weighted behavior cloning, allowing for more flexible policy learning while maintaining desirable behaviors. In the meta-testing phase, SELECT introduces a proximal context collection strategy that balances exploration and exploitation. This strategy gathers high-quality context, improving task inference and adaptation to unseen tasks. Experimental results show that SELECT significantly reduces the distributional shift, enhances the meta-policy's generalization, and outperforms state-of-the-art methods across various domains.
Paper
Full text
Offline Meta Reinforcement Learning with Weighted Policy Constraints and Proximal Context Collection
Semantic Scholar · Computer Science · 2025
Abstract
Offline meta-reinforcement learning (OMRL) encounters two key challenges: effectively learning the meta-policy from offline datasets and correctly inferring unseen tasks. Existing methods often address the first challenge by imposing policy constraints, but are limited by the suboptimal actions in offline datasets. For the second challenge, most focus on meta-training without enhancing task inference during meta-testing. To address these issues, we propose a novel method called weighted policy conStraints and proximal contExt coLlECtion sTrategy for OMRL (SELECT). During meta-training, we integrate policy constraints with weighted behavior cloning, allowing for more flexible policy learning while maintaining desirable behaviors. In the meta-testing phase, SELECT introduces a proximal context collection strategy that balances exploration and exploitation. This strategy gathers high-quality context, improving task inference and adaptation to unseen tasks. Experimental results show that SELECT significantly reduces the distributional shift, enhances the meta-policy's generalization, and outperforms state-of-the-art methods across various domains.