Post-hoc interpretability approaches have been proven to be powerful tools to\ngenerate explanations for the predictions made by a trained black-box model.\nHowever, they create the risk of having explanations that are a result of some\nartifacts learned by the model instead of actual knowledge from the data. This\npaper focuses on the case of counterfactual explanations and asks whether the\ngenerated instances can be justified, i.e. continuously connected to some\nground-truth data. We evaluate the risk of generating unjustified\ncounterfactual examples by investigating the local neighborhoods of instances\nwhose predictions are to be explained and show that this risk is quite high for\nseveral datasets. Furthermore, we show that most state of the art approaches do\nnot differentiate justified from unjustified counterfactual examples, leading\nto less useful explanations.\n
Paper
References (23)
Scroll for more · 11 remaining