Self-Supervised Multi-Agent Diversity with Nonparametric Entropy Maximization

Learning decentralized policies for agents has drawn increasing interest in recent works to solve the scalability issue that arises in Multi-Agent Reinforcement Learning (MARL), where all agents may share the parameters of a policy network to make action decisions. However, such parameter sharing can hinder efficient exploration, as some agents may learn similar behaviors. Unlike previous fully-supervised mutual information-based methods that encourages multi-agent diversity, in this paper, we propose a novel multi-agent exploration method called Contrastive Trajectory Entropy Maximization (CTEM). Our method adopts a non-parametric entropy estimator to maximize the entropy of trajectories of different agents in a self-supervised contrastive representation space, leading to diverse policies and sufficient exploration. Such an entropy estimator avoids complex density modeling and scales well in high-dimensional multi-agent environments. We deploy our method in MARL by introducing an intrinsic reward for agents to achieve entropy maximization. To demonstrate the effectiveness of our method, we conduct experiments on multiple challenging MARL benchmark tasks. Our method yields superior performance than existing state-of-the-art methods.

Paper

Full text

PDF

Self-Supervised Multi-Agent Diversity with Nonparametric Entropy Maximization

Semantic Scholar · Computer Science · 2025

Abstract

Learning decentralized policies for agents has drawn increasing interest in recent works to solve the scalability issue that arises in Multi-Agent Reinforcement Learning (MARL), where all agents may share the parameters of a policy network to make action decisions. However, such parameter sharing can hinder efficient exploration, as some agents may learn similar behaviors. Unlike previous fully-supervised mutual information-based methods that encourages multi-agent diversity, in this paper, we propose a novel multi-agent exploration method called Contrastive Trajectory Entropy Maximization (CTEM). Our method adopts a non-parametric entropy estimator to maximize the entropy of trajectories of different agents in a self-supervised contrastive representation space, leading to diverse policies and sufficient exploration. Such an entropy estimator avoids complex density modeling and scales well in high-dimensional multi-agent environments. We deploy our method in MARL by introducing an intrinsic reward for agents to achieve entropy maximization. To demonstrate the effectiveness of our method, we conduct experiments on multiple challenging MARL benchmark tasks. Our method yields superior performance than existing state-of-the-art methods.

Similar papers

© 2026 NYSGPT2525 LLC