In this paper, we consider the problem of learning independent fair policies in cooperative multi-agent reinforcement learning (MARL). Our objective is to design multiple policies simultaneously that optimize a welfare function for fairness. To achieve this objective, we propose a novel Fairness-Aware multi-agent Proximal Policy Optimization (FAPPO) algorithm, which enables each agent to learn its policy independently while optimizing a welfare function. Unlike standard approaches that focus on maximizing performance metrics such as rewards, FAPPO focuses on fairness in an independent learning setting, where each agent estimates its local value function. Furthermore, when inter-agent communication is allowed, we introduce an attention-based FAPPO (AT-FAPPO), which incorporates a self-attention mechanism to facilitate communication and coordination among agents. This variant allows agents to share relevant information during training, leading to more fair outcomes. To demonstrate the effectiveness of the proposed methods, we perform experiments in various environments and show that our approach outperforms existing methods both in terms of efficiency and equity.
Paper
Full text
Towards Fair and Efficient Policy Learning in Cooperative Multi-Agent Reinforcement Learning
Semantic Scholar · Computer Science · 2025
Abstract
In this paper, we consider the problem of learning independent fair policies in cooperative multi-agent reinforcement learning (MARL). Our objective is to design multiple policies simultaneously that optimize a welfare function for fairness. To achieve this objective, we propose a novel Fairness-Aware multi-agent Proximal Policy Optimization (FAPPO) algorithm, which enables each agent to learn its policy independently while optimizing a welfare function. Unlike standard approaches that focus on maximizing performance metrics such as rewards, FAPPO focuses on fairness in an independent learning setting, where each agent estimates its local value function. Furthermore, when inter-agent communication is allowed, we introduce an attention-based FAPPO (AT-FAPPO), which incorporates a self-attention mechanism to facilitate communication and coordination among agents. This variant allows agents to share relevant information during training, leading to more fair outcomes. To demonstrate the effectiveness of the proposed methods, we perform experiments in various environments and show that our approach outperforms existing methods both in terms of efficiency and equity.