Explainers in the Wild: Making Surrogate Explainers Robust to Distortions through Perception

Explaining the decisions of models is becoming pervasive in the image\nprocessing domain, whether it is by using post-hoc methods or by creating\ninherently interpretable models. While the widespread use of surrogate\nexplainers is a welcome addition to inspect and understand black-box models,\nassessing the robustness and reliability of the explanations is key for their\nsuccess. Additionally, whilst existing work in the explainability field\nproposes various strategies to address this problem, the challenges of working\nwith data in the wild is often overlooked. For instance, in image\nclassification, distortions to images can not only affect the predictions\nassigned by the model, but also the explanation. Given a clean and a distorted\nversion of an image, even if the prediction probabilities are similar, the\nexplanation may still be different. In this paper we propose a methodology to\nevaluate the effect of distortions in explanations by embedding perceptual\ndistances that tailor the neighbourhoods used to training surrogate explainers.\nWe also show that by operating in this way, we can make the explanations more\nrobust to distortions. We generate explanations for images in the Imagenet-C\ndataset and demonstrate how using a perceptual distances in the surrogate\nexplainer creates more coherent explanations for the distorted and reference\nimages.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC