Ethics-Aware Safe Reinforcement Learning for Rare-Event Risk Control in Interactive Urban Driving
Autonomous vehicles hold great promise for reducing traffic fatalities and improving transportation efficiency, yet their widespread adoption hinges on embedding credible and transparent ethical reasoning into routine and emergency maneuvers, particularly to protect vulnerable road users (VRUs) such as pedestrians and cyclists. Here, we present a hierarchical Safe Reinforcement Learning (Safe RL) framework that augments standard driving objectives with ethics-aware cost signals. At the decision level, a Safe RL agent is trained using a composite ethical risk cost, combining collision probability and harm severity, to generate high-level motion targets. To improve sample efficiency under rare high-risk events, we introduce a risk-sensitive prioritized experience replay scheme. We further propose Temporal Cost Aggregation (TCA), which propagates risk across decision steps and aligns learning with tail-risk measures such as Conditional Value-at-Risk (CVaR), mitigating single-step independence assumptions. At the execution level, polynomial trajectory generation coupled with Proportional-Integral-Derivative (PID) and Stanley controllers ensures smooth and feasible execution. We evaluate EthicAR in closed-loop simulations based on the Waymo Open Dataset across 75 real-world scenarios and five random seeds. The proposed method decreases collision rates by 20$\sim$45% compared to baseline methods, while maintaining task success rates and comfort metrics within 5$\sim$10% of baselines. This work provides a reproducible benchmark for Safe RL with explicitly ethics-aware objectives in human-mixed traffic scenarios. Our results highlight the potential of combining formal control theory and data-driven learning to advance ethically accountable autonomy that explicitly protects those most at risk in urban traffic environments.
Paper
References (34)
Scroll for more · 22 remaining