On the Approximation of Cooperative Heterogeneous Multi-Agent Reinforcement Learning (MARL) using Mean Field Control (MFC)
Mean field control (MFC) is an effective way to mitigate the curse of\ndimensionality of cooperative multi-agent reinforcement learning (MARL)\nproblems. This work considers a collection of $N_{\\mathrm{pop}}$ heterogeneous\nagents that can be segregated into $K$ classes such that the $k$-th class\ncontains $N_k$ homogeneous agents. We aim to prove approximation guarantees of\nthe MARL problem for this heterogeneous system by its corresponding MFC\nproblem. We consider three scenarios where the reward and transition dynamics\nof all agents are respectively taken to be functions of $(1)$ joint state and\naction distributions across all classes, $(2)$ individual distributions of each\nclass, and $(3)$ marginal distributions of the entire population. We show that,\nin these cases, the $K$-class MARL problem can be approximated by MFC with\nerrors given as\n$e_1=\\mathcal{O}(\\frac{\\sqrt{|\\mathcal{X}|}+\\sqrt{|\\mathcal{U}|}}{N_{\\mathrm{pop}}}\\sum_{k}\\sqrt{N_k})$,\n$e_2=\\mathcal{O}(\\left[\\sqrt{|\\mathcal{X}|}+\\sqrt{|\\mathcal{U}|}\\right]\\sum_{k}\\frac{1}{\\sqrt{N_k}})$\nand\n$e_3=\\mathcal{O}\\left(\\left[\\sqrt{|\\mathcal{X}|}+\\sqrt{|\\mathcal{U}|}\\right]\\left[\\frac{A}{N_{\\mathrm{pop}}}\\sum_{k\\in[K]}\\sqrt{N_k}+\\frac{B}{\\sqrt{N_{\\mathrm{pop}}}}\\right]\\right)$,\nrespectively, where $A, B$ are some constants and $|\\mathcal{X}|,|\\mathcal{U}|$\nare the sizes of state and action spaces of each agent. Finally, we design a\nNatural Policy Gradient (NPG) based algorithm that, in the three cases stated\nabove, can converge to an optimal MARL policy within $\\mathcal{O}(e_j)$ error\nwith a sample complexity of $\\mathcal{O}(e_j^{-3})$, $j\\in\\{1,2,3\\}$,\nrespectively.\n
Paper
References (40)
Scroll for more · 28 remaining