From Sparse to Dense: Toddler-inspired Reward Transition in Goal-Oriented Reinforcement Learning
Reinforcement learning (RL) agents face fundamental challenges in balancing exploration and exploitation, particularly when sparse or dense rewards bias learning toward suboptimal behaviors. Biological systems, notably human toddlers, naturally navigate this balance through developmental transitions from free exploration to goal-directed behavior. Inspired by this progression, we propose the sparse-to-dense (S2D) reward transition for goal-oriented RL tasks. Our approach transitions from sparse rewards encouraging broad exploration to dense rewards guiding efficient exploitation, while preserving optimal policies via potential-based reward shaping throughout. Through extensive experiments on robotic manipulation and egocentric 3-D navigation tasks, we demonstrate that S2D transitions significantly enhance sample efficiency and generalization compared to static reward schemes and intrinsic motivation baselines. Using a novel cross-density visualizer, we reveal that S2D transitions smooth the policy loss landscape, resulting in wider minima associated with improved generalization. Furthermore, reinterpreting Tolman’s maze experiments in modern RL contexts, we show that early sparse-reward exploration establishes robust initial parameters analogous to cognitive maps, facilitating stable learning under subsequent dense rewards. Our findings bridge developmental psychology and machine learning, offering a principled framework for designing adaptive reward structures in complex RL environments.