Designing incentive mechanisms for multi-agent systems in stochastic and dynamic environments is a critical challenge, as system outcomes emerge from the complex interplay of agent learning and environmental uncertainty. Existing principal–multi-agent contract design methods often assume static settings or ignore learning dynamics, limiting their applicability in multi-agent reinforcement learning (MARL). Furthermore, the contract design space is highly constrained by feasibility requirements, such as individual rationality and incentive compatibility, making it difficult to explore. We introduce the principal-MARL contract design problem, where a principal optimizes both recruitment and incentive contracts evaluated via MARL. To address this problem, we propose Constrained Pareto Maximum Entropy Search (cPMES), a multi-objective Bayesian optimization framework that treats feasibility as an explicit objective and selects designs based on information gain over the Pareto front. Experiments in social dilemma environments demonstrate that cPMES efficiently identifies feasible, high-performing contracts, significantly improving coordination and system-level rewards.
Paper
References (25)
Scroll for more · 13 remaining