Neural Tree Expansion for Multi-Robot Planning in Non-Cooperative Environments

We present a self-improving, Neural Tree Expansion (NTE) method for\nmulti-robot online planning in non-cooperative environments, where each robot\nattempts to maximize its cumulative reward while interacting with other\nself-interested robots. Our algorithm adapts the centralized, perfect\ninformation, discrete-action space method from AlphaZero to a decentralized,\npartial information, continuous action space setting for multi-robot\napplications. Our method has three interacting components: (i) a centralized,\nperfect-information "expert" Monte Carlo Tree Search (MCTS) with large\ncomputation resources that provides expert demonstrations, (ii) a\ndecentralized, partial-information "learner" MCTS with small computation\nresources that runs in real-time and provides self-play examples, and (iii)\npolicy & value neural networks that are trained with the expert demonstrations\nand bias both the expert and the learner tree growth. Our numerical experiments\ndemonstrate Neural Tree Expansion's computational advantage by finding better\nsolutions than a MCTS with 20 times more resources. The resulting policies are\ndynamically sophisticated, demonstrate coordination between robots, and play\nthe Reach-Target-Avoid differential game significantly better than the\nstate-of-the-art control-theoretic baseline for multi-robot, double-integrator\nsystems. Our hardware experiments on an aerial swarm demonstrate the\ncomputational advantage of Neural Tree Expansion, enabling online planning at\n20Hz with effective policies in complex scenarios.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC