Learning Self-Game-Play Agents for Combinatorial Optimization Problems

Recent progress in reinforcement learning (RL) using self-game-play has shown remarkable performance on several board games as well as video games (e.g., Atari games and Dota2). DeepMind researchers have already implemented model-free RL to play Go and Chess at a superhuman level using neural Monte-Carlo-Tree-Search (neural MCTS). Therefore, it is plausible to consider that RL, starting from zero knowledge, can be applied to other problems which can be converted into games. We try to leverage the computational power of neural MCTS to solve a class of combinatorial optimization problems. Following the idea of Hintikka's Game-Theoretical Semantics, we propose the Zermelo Gamification (ZG) to transform specific combinatorial optimization problems into Zermelo games whose winning strategies correspond to the solutions of the original optimization problem. The ZG also provides a specially designed neural MCTS. We use a combinatorial planning problem for which the ground-truth policy is efficiently computable to demonstrate that ZG is promising.

Paper

Similar papers

© 2026 NYSGPT2525 LLC