Dynamic Incentivized Cooperation under Changing Rewards

Many real-world multi-agent systems are characterized by two simultaneous challenges: strategic tension in social dilemmas and non-stationary reward signals. While peer incentivization (PI) has emerged as a decentralized mechanism to promote cooperation in multi-agent reinforcement learning (MARL), existing approaches typically rely on fixed or externally scaled incentive magnitudes. When environmental rewards change, due to scaling, shifting, or drift, the relative strength between rewards and incentives can become misaligned, which destabilizes cooperation even when the underlying strategic structure remains unchanged. We analyze this structural sensitivity and argue that reward normalization preserves gradient invariance but does not resolve incentive misalignment in social dilemmas. We then introduce Dynamic Reward Incentives for Variable Exchange (DRIVE), a reciprocal shaping mechanism that exchanges reward differences rather than fixed magnitudes. Because these differences are expressed in reward units, they scale proportionally under affine reward changes, preserving the relative influence of environmental rewards and incentives.

Paper

Similar papers

© 2026 NYSGPT2525 LLC