Proximal Policy Optimization in Autonomous Driving: A Systematic Review of Methods, Imitation Learning, and Evaluation Practices
The development of robust decision-making policies for road-based autonomous vehicles (AV) remains a critical research challenge. While Reinforcement Learning (RL) and Imitation Learning (IL) show promise, the research landscape remains fragmented, particularly regarding Proximal Policy Optimization (PPO). This paper presents a systematic literature review synthesizing PPO applications in autonomous driving. Following the Kitchenham methodology and SEGRESS guidelines, we searched major digital libraries for studies published between 2021 and 2025. From 108 initial records, a rigorous selection process yielded 26 primary studies. Our analysis reveals that standard PPO dominates (73.1%), with modified variants accounting for 23.1%. CARLA serves as the primary evaluation platform (50.0%), with urban driving (42.3%) and lane changing (30.8%) being the most common tasks. IL integration employed diverse approaches including Behavioral Cloning, GAIL, and AIRL. A significant finding is evaluation fragmentation, with 14 unique metrics identified, though collision rate (57.7%) and success rate (30.8%) were most prevalent. Quality assessment showed 46.2% of studies achieved high methodological quality, while transparency emerged as the weakest criterion. Our findings underscore the need for standardized benchmarks, sim-to-real transfer methods, and improved reporting transparency.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex