Strategically Efficient Exploration in Competitive Multi-agent Reinforcement Learning

High sample complexity remains a barrier to the application of reinforcement\nlearning (RL), particularly in multi-agent systems. A large body of work has\ndemonstrated that exploration mechanisms based on the principle of optimism\nunder uncertainty can significantly improve the sample efficiency of RL in\nsingle agent tasks. This work seeks to understand the role of optimistic\nexploration in non-cooperative multi-agent settings. We will show that, in\nzero-sum games, optimistic exploration can cause the learner to waste time\nsampling parts of the state space that are irrelevant to strategic play, as\nthey can only be reached through cooperation between both players. To address\nthis issue, we introduce a formal notion of strategically efficient exploration\nin Markov games, and use this to develop two strategically efficient learning\nalgorithms for finite Markov games. We demonstrate that these methods can be\nsignificantly more sample efficient than their optimistic counterparts.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC