Temporal-Spatial Causal Interpretations for Vision-Based Reinforcement Learning

Deep reinforcement learning (RL) agents are becoming increasingly proficient\nin a range of complex control tasks. However, the agent's behavior is usually\ndifficult to interpret due to the introduction of black-box function, making it\ndifficult to acquire the trust of users. Although there have been some\ninteresting interpretation methods for vision-based RL, most of them cannot\nuncover temporal causal information, raising questions about their reliability.\nTo address this problem, we present a temporal-spatial causal interpretation\n(TSCI) model to understand the agent's long-term behavior, which is essential\nfor sequential decision-making. TSCI model builds on the formulation of\ntemporal causality, which reflects the temporal causal relations between\nsequential observations and decisions of RL agent. Then a separate causal\ndiscovery network is employed to identify temporal-spatial causal features,\nwhich are constrained to satisfy the temporal causality. TSCI model is\napplicable to recurrent agents and can be used to discover causal features with\nhigh efficiency once trained. The empirical results show that TSCI model can\nproduce high-resolution and sharp attention masks to highlight task-relevant\ntemporal-spatial information that constitutes most evidence about how\nvision-based RL agents make sequential decisions. In addition, we further\ndemonstrate that our method is able to provide valuable causal interpretations\nfor vision-based RL agents from the temporal perspective.\n

Paper

References (70)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC