Model-agnostic and Scalable Counterfactual Explanations via Reinforcement Learning

Counterfactual instances are a powerful tool to obtain valuable insights into\nautomated decision processes, describing the necessary minimal changes in the\ninput space to alter the prediction towards a desired target. Most previous\napproaches require a separate, computationally expensive optimization procedure\nper instance, making them impractical for both large amounts of data and\nhigh-dimensional data. Moreover, these methods are often restricted to certain\nsubclasses of machine learning models (e.g. differentiable or tree-based\nmodels). In this work, we propose a deep reinforcement learning approach that\ntransforms the optimization procedure into an end-to-end learnable process,\nallowing us to generate batches of counterfactual instances in a single forward\npass. Our experiments on real-world data show that our method i) is\nmodel-agnostic (does not assume differentiability), relying only on feedback\nfrom model predictions; ii) allows for generating target-conditional\ncounterfactual instances; iii) allows for flexible feature range constraints\nfor numerical and categorical attributes, including the immutability of\nprotected features (e.g. gender, race); iv) is easily extended to other data\nmodalities such as images.\n

Paper

References (41)

Scroll for more · 29 remaining

Similar papers

© 2026 NYSGPT2525 LLC