DisCo RL: Distribution-Conditioned Reinforcement Learning for General-Purpose Policies

Can we use reinforcement learning to learn general-purpose policies that can\nperform a wide range of different tasks, resulting in flexible and reusable\nskills? Contextual policies provide this capability in principle, but the\nrepresentation of the context determines the degree of generalization and\nexpressivity. Categorical contexts preclude generalization to entirely new\ntasks. Goal-conditioned policies may enable some generalization, but cannot\ncapture all tasks that might be desired. In this paper, we propose goal\ndistributions as a general and broadly applicable task representation suitable\nfor contextual policies. Goal distributions are general in the sense that they\ncan represent any state-based reward function when equipped with an appropriate\ndistribution class, while the particular choice of distribution class allows us\nto trade off expressivity and learnability. We develop an off-policy algorithm\ncalled distribution-conditioned reinforcement learning (DisCo RL) to\nefficiently learn these policies. We evaluate DisCo RL on a variety of robot\nmanipulation tasks and find that it significantly outperforms prior methods on\ntasks that require generalization to new goal distributions.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC