Influence-Based Reinforcement Learning for Intrinsically-Motivated Agents

Discovering successful coordinated behaviors is a central challenge in\nMulti-Agent Reinforcement Learning (MARL) since it requires exploring a joint\naction space that grows exponentially with the number of agents. In this paper,\nwe propose a mechanism for achieving sufficient exploration and coordination in\na team of agents. Specifically, agents are rewarded for contributing to a more\ndiversified team behavior by employing proper intrinsic motivation functions.\nTo learn meaningful coordination protocols, we structure agents' interactions\nby introducing a novel framework, where at each timestep, an agent simulates\ncounterfactual rollouts of its policy and, through a sequence of computations,\nassesses the gap between other agents' current behaviors and their targets.\nActions that minimize the gap are considered highly influential and are\nrewarded. We evaluate our approach on a set of challenging tasks with sparse\nrewards and partial observability that require learning complex cooperative\nstrategies under a proper exploration scheme, such as the StarCraft Multi-Agent\nChallenge. Our methods show significantly improved performances over different\nbaselines across all tasks.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC