RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments

Exploration in sparse reward environments remains one of the key challenges\nof model-free reinforcement learning. Instead of solely relying on extrinsic\nrewards provided by the environment, many state-of-the-art methods use\nintrinsic rewards to encourage exploration. However, we show that existing\nmethods fall short in procedurally-generated environments where an agent is\nunlikely to visit a state more than once. We propose a novel type of intrinsic\nreward which encourages the agent to take actions that lead to significant\nchanges in its learned state representation. We evaluate our method on multiple\nchallenging procedurally-generated tasks in MiniGrid, as well as on tasks with\nhigh-dimensional observations used in prior work. Our experiments demonstrate\nthat this approach is more sample efficient than existing exploration methods,\nparticularly for procedurally-generated MiniGrid environments. Furthermore, we\nanalyze the learned behavior as well as the intrinsic reward received by our\nagent. In contrast to previous approaches, our intrinsic reward does not\ndiminish during the course of training and it rewards the agent substantially\nmore for interacting with objects that it can control.\n

Paper

References (67)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC