Zero-Shot Coordination in Ad Hoc Teams with Generalized Policy Improvement and Difference Rewards
Real-world multi-agent systems may require ad hoc teaming, where an agent must coordinate with previously unseen teammates in a zero-shot manner. Prior work either selects a pretrained policy based on an inferred model of the new teammates or pretrains a single policy that is robust to potential teammates. Instead, we propose to dynamically leverage all pretrained policies through two key ideas---generalized policy improvement and difference rewards---for efficient and effective knowledge transfer between different teams. We empirically demonstrate that our algorithm, Generalized Policy improvement for Ad hoc Teaming (GPAT), successfully enables zero-shot transfer to new teams in three simulated environments and demonstrate our algorithm in a real-world multi-robot setting.
Paper
References (38)
Scroll for more · 26 remaining