$α$-Rank: Multi-Agent Evaluation by Evolution

We introduce $\\alpha$-Rank, a principled evolutionary dynamics methodology\nfor the evaluation and ranking of agents in large-scale multi-agent\ninteractions, grounded in a novel dynamical game-theoretic solution concept\ncalled Markov-Conley chains (MCCs). The approach leverages continuous- and\ndiscrete-time evolutionary dynamical systems applied to empirical games, and\nscales tractably in the number of agents, the type of interactions, and the\ntype of empirical games (symmetric and asymmetric). Current models are\nfundamentally limited in one or more of these dimensions and are not guaranteed\nto converge to the desired game-theoretic solution concept (typically the Nash\nequilibrium). $\\alpha$-Rank provides a ranking over the set of agents under\nevaluation and provides insights into their strengths, weaknesses, and\nlong-term dynamics. This is a consequence of the links we establish to the MCC\nsolution concept when the underlying evolutionary model's ranking-intensity\nparameter, $\\alpha$, is chosen to be large, which exactly forms the basis of\n$\\alpha$-Rank. In contrast to the Nash equilibrium, which is a static concept\nbased on fixed points, MCCs are a dynamical solution concept based on the\nMarkov chain formalism, Conley's Fundamental Theorem of Dynamical Systems, and\nthe core ingredients of dynamical systems: fixed points, recurrent sets,\nperiodic orbits, and limit cycles. $\\alpha$-Rank runs in polynomial time with\nrespect to the total number of pure strategy profiles, whereas computing a Nash\nequilibrium for a general-sum game is known to be intractable. We introduce\nproofs that not only provide a unifying perspective of existing continuous- and\ndiscrete-time evolutionary evaluation models, but also reveal the formal\nunderpinnings of the $\\alpha$-Rank methodology. We empirically validate the\nmethod in several domains including AlphaGo, AlphaZero, MuJoCo Soccer, and\nPoker.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC