Reducing Variance Caused by Communication in Decentralized Multi-agent Deep Reinforcement Learning
In decentralized multi-agent deep reinforcement learning (MADRL), communication can help agents to gain a better understanding of the environment to better coordinate their behaviors. Nevertheless, communication may involve uncertainty, which potentially introduces variance to the learning of decentralized agents. In this extended abstract, we report on our research that focuses on a specific decentralized MADRL setting with communication and a theoretical analysis to study the variance caused by communication in policy gradients. We argue for modular techniques to reduce the variance in policy gradients during training. We show a pseudo algorithm to illustrate the integration of the modular techniques into existing decentralized MADRL with communication methods.