One major barrier to applications of deep Reinforcement Learning (RL) both\ninside and outside of games is the lack of explainability. In this paper, we\ndescribe a lightweight and effective method to derive explanations for deep RL\nagents, which we evaluate in the Atari domain. Our method relies on a\ntransformation of the pixel-based input of the RL agent to an interpretable,\npercept-like input representation. We then train a surrogate model, which is\nitself interpretable, to replicate the behavior of the target, deep RL agent.\nOur experiments demonstrate that we can learn an effective surrogate that\naccurately approximates the underlying decision making of a target agent on a\nsuite of Atari games.\n
Paper
References (39)
Scroll for more · 27 remaining