Contrasting Centralized and Decentralized Critics in Multi-Agent Reinforcement Learning

Centralized Training for Decentralized Execution, where agents are trained\noffline using centralized information but execute in a decentralized manner\nonline, has gained popularity in the multi-agent reinforcement learning\ncommunity. In particular, actor-critic methods with a centralized critic and\ndecentralized actors are a common instance of this idea. However, the\nimplications of using a centralized critic in this context are not fully\ndiscussed and understood even though it is the standard choice of many\nalgorithms. We therefore formally analyze centralized and decentralized critic\napproaches, providing a deeper understanding of the implications of critic\nchoice. Because our theory makes unrealistic assumptions, we also empirically\ncompare the centralized and decentralized critic methods over a wide set of\nenvironments to validate our theories and to provide practical advice. We show\nthat there exist misconceptions regarding centralized critics in the current\nliterature and show that the centralized critic design is not strictly\nbeneficial, but rather both centralized and decentralized critics have\ndifferent pros and cons that should be taken into account by algorithm\ndesigners.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC