A Meta Multi-agent Reinforcement Learning Algorithm for Multi-intersection Traffic Signal Control

Multi-agent deep reinforcement learning (MDRL) has made remarkable progress in the multi-intersection traffic signal control, and meta-learning further promotes its learning capability. However, most existing MDRL algorithms seem to have some drawbacks. 1) These algorithms only employ the information at each step (i.e., short-term information) to train the models, which may lead to induce the non-optimal traffic-signal policies due to the long-term information (e.g., tasks of each agent) ignored. 2) These algorithms based on meta-learning cannot effectively tackle diversity among the tasks due to the shared parameters describing the ‘average’ source tasks. Motivated by the above observations, in this paper, we propose an algorithm, referred to as Meta Multi-agent Advantage Actor-critic (ME-MA2C) algorithm, for multi-intersection traffic signal control. The proposed ME-MA2C algorithm has two components: 1) It conducts meta-learning (ME) using a proposed task-neighbor encoder, which is a meta-learning algorithm. ME algorithm encodes both the short-term information and the long-term information to learn the meta embedding and meta-knowledge, which helps to induce optimal traffic-signal policies. 2) It also conducts policy learning using a Multi-agent Advantage Actor-Critic (MA2C), which is a decentralized multi-agent framework. MA2C framework makes use of the learned meta embeddings and meta-knowledge, and optimizes the whole algorithm to derive transferable traffic-signal policies. Experimental results on different datasets illustrate that ME-MA2C algorithm outperforms the state-of-the-art algorithms in terms of multiple metrics, which seems to be a new promising algorithm for multi-intersection traffic signal control.

Paper

Full text

PDF

A Meta Multi-agent Reinforcement Learning Algorithm for Multi-intersection Traffic Signal Control

Semantic Scholar · Computer Science · 2021

Abstract

Multi-agent deep reinforcement learning (MDRL) has made remarkable progress in the multi-intersection traffic signal control, and meta-learning further promotes its learning capability. However, most existing MDRL algorithms seem to have some drawbacks. 1) These algorithms only employ the information at each step (i.e., short-term information) to train the models, which may lead to induce the non-optimal traffic-signal policies due to the long-term information (e.g., tasks of each agent) ignored. 2) These algorithms based on meta-learning cannot effectively tackle diversity among the tasks due to the shared parameters describing the ‘average’ source tasks. Motivated by the above observations, in this paper, we propose an algorithm, referred to as Meta Multi-agent Advantage Actor-critic (ME-MA2C) algorithm, for multi-intersection traffic signal control. The proposed ME-MA2C algorithm has two components: 1) It conducts meta-learning (ME) using a proposed task-neighbor encoder, which is a meta-learning algorithm. ME algorithm encodes both the short-term information and the long-term information to learn the meta embedding and meta-knowledge, which helps to induce optimal traffic-signal policies. 2) It also conducts policy learning using a Multi-agent Advantage Actor-Critic (MA2C), which is a decentralized multi-agent framework. MA2C framework makes use of the learned meta embeddings and meta-knowledge, and optimizes the whole algorithm to derive transferable traffic-signal policies. Experimental results on different datasets illustrate that ME-MA2C algorithm outperforms the state-of-the-art algorithms in terms of multiple metrics, which seems to be a new promising algorithm for multi-intersection traffic signal control.

Similar papers

© 2026 NYSGPT2525 LLC