Model-based Reinforcement Learning for Decentralized Multiagent Rendezvous

Collaboration requires agents to align their goals on the fly. Underlying the\nhuman ability to align goals with other agents is their ability to predict the\nintentions of others and actively update their own plans. We propose\nhierarchical predictive planning (HPP), a model-based reinforcement learning\nmethod for decentralized multiagent rendezvous. Starting with pretrained,\nsingle-agent point to point navigation policies and using noisy,\nhigh-dimensional sensor inputs like lidar, we first learn via self-supervision\nmotion predictions of all agents on the team. Next, HPP uses the prediction\nmodels to propose and evaluate navigation subgoals for completing the\nrendezvous task without explicit communication among agents. We evaluate HPP in\na suite of unseen environments, with increasing complexity and numbers of\nobstacles. We show that HPP outperforms alternative reinforcement learning,\npath planning, and heuristic-based baselines on challenging, unseen\nenvironments. Experiments in the real world demonstrate successful transfer of\nthe prediction models from sim to real world without any additional\nfine-tuning. Altogether, HPP removes the need for a centralized operator in\nmultiagent systems by combining model-based RL and inference methods, enabling\nagents to dynamically align plans.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC