Autonomous vehicles need social awareness to find optima in multi-agent reinforcement learning routing games

Previous work has shown that when multiple selfish Autonomous Vehicles (AVs) simultaneously learn optimal routing strategies using Multi-Agent Reinforcement Learning (MARL), they may require a significant amount of time to converge to the optimal solution, equivalent to years of real-world commuting. We demonstrate that moving beyond the selfish component in the reward significantly relieves this issue. In particular, we introduce a reward signal based on the marginal cost matrix, which quantifies the impact of each individual action (route-choice) on the system (total travel time). This formulation reduces training time and improves convergence reliability. Experiments on both a toy network and the real-world Saint-Arnoult network show that the proposed reward improves individual and system travel times over the selfish reward baseline, and in the toy network, enables agents to reach the optimal solution faster, indicating that incorporating social awareness (i.e., including marginal costs in routing decisions) can enhance both system-wide and individual outcomes in future urban systems with AVs.

Paper

References (41)

Scroll for more · 29 remaining

Similar papers

© 2026 NYSGPT2525 LLC