Plausible Counterfactuals: Auditing Deep Learning Classifiers with Realistic Adversarial Examples
The last decade has witnessed the proliferation of Deep Learning models in\nmany applications, achieving unrivaled levels of predictive performance.\nUnfortunately, the black-box nature of Deep Learning models has posed\nunanswered questions about what they learn from data. Certain application\nscenarios have highlighted the importance of assessing the bounds under which\nDeep Learning models operate, a problem addressed by using assorted approaches\naimed at audiences from different domains. However, as the focus of the\napplication is placed more on non-expert users, it results mandatory to provide\nthe means for him/her to trust the model, just like a human gets familiar with\na system or process: by understanding the hypothetical circumstances under\nwhich it fails. This is indeed the angular stone for this research work: to\nundertake an adversarial analysis of a Deep Learning model. The proposed\nframework constructs counterfactual examples by ensuring their plausibility,\ne.g. there is a reasonable probability that a human could generate them without\nresorting to a computer program. Therefore, this work must be regarded as\nvaluable auditing exercise of the usable bounds a certain model is constrained\nwithin, thereby allowing for a much greater understanding of the capabilities\nand pitfalls of a model used in a real application. To this end, a Generative\nAdversarial Network (GAN) and multi-objective heuristics are used to furnish a\nplausible attack to the audited model, efficiently trading between the\nconfusion of this model, the intensity and plausibility of the generated\ncounterfactual. Its utility is showcased within a human face classification\ntask, unveiling the enormous potential of the proposed framework.\n