A Novel Automated Curriculum Strategy to Solve Hard Sokoban Planning Instances

In recent years, we have witnessed tremendous progress in deep reinforcement\nlearning (RL) for tasks such as Go, Chess, video games, and robot control.\nNevertheless, other combinatorial domains, such as AI planning, still pose\nconsiderable challenges for RL approaches. The key difficulty in those domains\nis that a positive reward signal becomes {\\em exponentially rare} as the\nminimal solution length increases. So, an RL approach loses its training\nsignal. There has been promising recent progress by using a curriculum-driven\nlearning approach that is designed to solve a single hard instance. We present\na novel {\\em automated} curriculum approach that dynamically selects from a\npool of unlabeled training instances of varying task complexity guided by our\n{\\em difficulty quantum momentum} strategy. We show how the smoothness of the\ntask hardness impacts the final learning results. In particular, as the size of\nthe instance pool increases, the ``hardness gap'' decreases, which facilitates\na smoother automated curriculum based learning process. Our automated\ncurriculum approach dramatically improves upon the previous approaches. We show\nour results on Sokoban, which is a traditional PSPACE-complete planning problem\nand presents a great challenge even for specialized solvers. Our RL agent can\nsolve hard instances that are far out of reach for any previous\nstate-of-the-art Sokoban solver. In particular, our approach can uncover plans\nthat require hundreds of steps, while the best previous search methods would\ntake many years of computing time to solve such instances. In addition, we show\nthat we can further boost the RL performance with an intricate coupling of our\nautomated curriculum approach with a curiosity-driven search strategy and a\ngraph neural net representation.\n

Paper

References (26)

Scroll for more · 14 remaining

Similar papers

© 2026 NYSGPT2525 LLC