How to Learn from Risk: Explicit Risk-Utility Reinforcement Learning for Efficient and Safe Driving Strategies
Autonomous driving has the potential to revolutionize mobility and is hence\nan active area of research. In practice, the behavior of autonomous vehicles\nmust be acceptable, i.e., efficient, safe, and interpretable. While vanilla\nreinforcement learning (RL) finds performant behavioral strategies, they are\noften unsafe and uninterpretable. Safety is introduced through Safe RL\napproaches, but they still mostly remain uninterpretable as the learned\nbehaviour is jointly optimized for safety and performance without modeling them\nseparately. Interpretable machine learning is rarely applied to RL. This paper\nproposes SafeDQN, which allows to make the behavior of autonomous vehicles safe\nand interpretable while still being efficient. SafeDQN offers an\nunderstandable, semantic trade-off between the expected risk and the utility of\nactions while being algorithmically transparent. We show that SafeDQN finds\ninterpretable and safe driving policies for a variety of scenarios and\ndemonstrate how state-of-the-art saliency techniques can help to assess both\nrisk and utility.\n
Paper
References (37)
Scroll for more · 25 remaining