Consensus Multiplicative Weights Update: Learning to Learn using Projector-based Game Signatures
Cheung and Piliouras (2020) recently showed that two variants of the\nMultiplicative Weights Update method - OMWU and MWU - display opposite\nconvergence properties depending on whether the game is zero-sum or\ncooperative. Inspired by this work and the recent literature on learning to\noptimize for single functions, we introduce a new framework for learning\nlast-iterate convergence to Nash Equilibria in games, where the update rule's\ncoefficients (learning rates) along a trajectory are learnt by a reinforcement\nlearning policy that is conditioned on the nature of the game: \\textit{the game\nsignature}. We construct the latter using a new decomposition of two-player\ngames into eight components corresponding to commutative projection operators,\ngeneralizing and unifying recent game concepts studied in the literature. We\ncompare the performance of various update rules when their coefficients are\nlearnt, and show that the RL policy is able to exploit the game signature\nacross a wide range of game types. In doing so, we introduce CMWU, a new\nalgorithm that extends consensus optimization to the constrained case, has\nlocal convergence guarantees for zero-sum bimatrix games, and show that it\nenjoys competitive performance on both zero-sum games with constant\ncoefficients and across a spectrum of games when its coefficients are learnt.\n