Greedy-based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning
Due to the representation limitation of the joint Q value function,\nmulti-agent reinforcement learning methods with linear value decomposition\n(LVD) or monotonic value decomposition (MVD) suffer from relative\novergeneralization. As a result, they can not ensure optimal consistency (i.e.,\nthe correspondence between individual greedy actions and the maximal true Q\nvalue). In this paper, we derive the expression of the joint Q value function\nof LVD and MVD. According to the expression, we draw a transition diagram,\nwhere each self-transition node (STN) is a possible convergence. To ensure\noptimal consistency, the optimal node is required to be the unique STN.\nTherefore, we propose the greedy-based value representation (GVR), which turns\nthe optimal node into an STN via inferior target shaping and further eliminates\nthe non-optimal STNs via superior experience replay. In addition, GVR achieves\nan adaptive trade-off between optimality and stability. Our method outperforms\nstate-of-the-art baselines in experiments on various benchmarks. Theoretical\nproofs and empirical results on matrix games demonstrate that GVR ensures\noptimal consistency under sufficient exploration.\n
Paper
References (25)
Scroll for more · 13 remaining