Exploring Counterfactual Explanations Through the Lens of Adversarial Examples: A Theoretical and Empirical Analysis
As machine learning (ML) models become more widely deployed in high-stakes\napplications, counterfactual explanations have emerged as key tools for\nproviding actionable model explanations in practice. Despite the growing\npopularity of counterfactual explanations, a deeper understanding of these\nexplanations is still lacking. In this work, we systematically analyze\ncounterfactual explanations through the lens of adversarial examples. We do so\nby formalizing the similarities between popular counterfactual explanation and\nadversarial example generation methods identifying conditions when they are\nequivalent. We then derive the upper bounds on the distances between the\nsolutions output by counterfactual explanation and adversarial example\ngeneration methods, which we validate on several real-world data sets. By\nestablishing these theoretical and empirical similarities between\ncounterfactual explanations and adversarial examples, our work raises\nfundamental questions about the design and development of existing\ncounterfactual explanation algorithms.\n