TripleTree: A Versatile Interpretable Representation of Black Box Agents and their Environments

In explainable artificial intelligence, there is increasing interest in\nunderstanding the behaviour of autonomous agents to build trust and validate\nperformance. Modern agent architectures, such as those trained by deep\nreinforcement learning, are currently so lacking in interpretable structure as\nto effectively be black boxes, but insights may still be gained from an\nexternal, behaviourist perspective. Inspired by conceptual spaces theory, we\nsuggest that a versatile first step towards general understanding is to\ndiscretise the state space into convex regions, jointly capturing similarities\nover the agent's action, value function and temporal dynamics within a dataset\nof observations. We create such a representation using a novel variant of the\nCART decision tree algorithm, and demonstrate how it facilitates practical\nunderstanding of black box agents through prediction, visualisation and\nrule-based explanation.\n

Paper

References (30)

Scroll for more · 18 remaining

Similar papers

© 2026 NYSGPT2525 LLC