Learning Long-Term Reward Redistribution via Randomized Return Decomposition

Many practical applications of reinforcement learning require agents to learn\nfrom sparse and delayed rewards. It challenges the ability of agents to\nattribute their actions to future outcomes. In this paper, we consider the\nproblem formulation of episodic reinforcement learning with trajectory\nfeedback. It refers to an extreme delay of reward signals, in which the agent\ncan only obtain one reward signal at the end of each trajectory. A popular\nparadigm for this problem setting is learning with a designed auxiliary dense\nreward function, namely proxy reward, instead of sparse environmental signals.\nBased on this framework, this paper proposes a novel reward redistribution\nalgorithm, randomized return decomposition (RRD), to learn a proxy reward\nfunction for episodic reinforcement learning. We establish a surrogate problem\nby Monte-Carlo sampling that scales up least-squares-based reward\nredistribution to long-horizon problems. We analyze our surrogate loss function\nby connection with existing methods in the literature, which illustrates the\nalgorithmic properties of our approach. In experiments, we extensively evaluate\nour proposed method on a variety of benchmark tasks with episodic rewards and\ndemonstrate substantial improvement over baseline algorithms.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC