An Analysis of Connections Between Regret Minimization and Actor Critic Methods in Cooperative Settings
Counterfactual Multi-agent Policy Gradients (COMA) is a popular algorithm for learning in cooperative multi-agent reinforcement learning settings. COMA computes difference rewards to solve the multiagent credit assignment problem by providing a local learning signal for each agent. Similar to other popular Cooperative multiagent RL (MARL) algorithms, there is a lack of theoretical justification for COMA's empirical success and specific way of doing credit assignment using difference rewards. We provide such a justification by connecting COMA's update rule to regret minimization. We then use this connection to improve COMA's performance by replacing usual softmax update with Neural Replicator Dynamics update from regret minimization literature. Experimental results on Starcraft II maps show the relevance of these theoretical insights for the performance of COMA in practice.
Paper
Full text
An Analysis of Connections Between Regret Minimization and Actor Critic Methods in Cooperative Settings
Semantic Scholar · Computer Science · 2023
Abstract
Counterfactual Multi-agent Policy Gradients (COMA) is a popular algorithm for learning in cooperative multi-agent reinforcement learning settings. COMA computes difference rewards to solve the multiagent credit assignment problem by providing a local learning signal for each agent. Similar to other popular Cooperative multiagent RL (MARL) algorithms, there is a lack of theoretical justification for COMA's empirical success and specific way of doing credit assignment using difference rewards. We provide such a justification by connecting COMA's update rule to regret minimization. We then use this connection to improve COMA's performance by replacing usual softmax update with Neural Replicator Dynamics update from regret minimization literature. Experimental results on Starcraft II maps show the relevance of these theoretical insights for the performance of COMA in practice.