Guided Exploration with Proximal Policy Optimization using a Single Demonstration

Solving sparse reward tasks through exploration is one of the major\nchallenges in deep reinforcement learning, especially in three-dimensional,\npartially-observable environments. Critically, the algorithm proposed in this\narticle uses a single human demonstration to solve hard-exploration problems.\nWe train an agent on a combination of demonstrations and own experience to\nsolve problems with variable initial conditions. We adapt this idea and\nintegrate it with the proximal policy optimization (PPO). The agent is able to\nincrease its performance and to tackle harder problems by replaying its own\npast trajectories prioritizing them based on the obtained reward and the\nmaximum value of the trajectory. We compare different variations of this\nalgorithm to behavioral cloning on a set of hard-exploration tasks in the\nAnimal-AI Olympics environment. To the best of our knowledge, learning a task\nin a three-dimensional environment with comparable difficulty has never been\nconsidered before using only one human demonstration.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC