Conservative Q-Improvement: Reinforcement Learning for an Interpretable Decision-Tree Policy

There is a growing desire in the field of reinforcement learning (and machine\nlearning in general) to move from black-box models toward more "interpretable\nAI." We improve interpretability of reinforcement learning by increasing the\nutility of decision tree policies learned via reinforcement learning. These\npolicies consist of a decision tree over the state space, which requires fewer\nparameters to express than traditional policy representations. Existing methods\nfor creating decision tree policies via reinforcement learning focus on\naccurately representing an action-value function during training, but this\nleads to much larger trees than would otherwise be required. To address this\nshortcoming, we propose a novel algorithm which only increases tree size when\nthe estimated discounted future reward of the overall policy would increase by\na sufficient amount. Through evaluation in a simulated environment, we show\nthat its performance is comparable or superior to traditional tree-based\napproaches and that it yields a more succinct policy. Additionally, we discuss\ntuning parameters to control the tradeoff between optimizing for smaller tree\nsize or for overall reward.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC