A number of techniques have been proposed to explain a machine learning\nmodel's prediction by attributing it to the corresponding input features.\nPopular among these are techniques that apply the Shapley value method from\ncooperative game theory. While existing papers focus on the axiomatic\nmotivation of Shapley values, and efficient techniques for computing them, they\noffer little justification for the game formulations used, and do not address\nthe uncertainty implicit in their methods' outputs. For instance, the popular\nSHAP algorithm's formulation may give substantial attributions to features that\nplay no role in the model. In this work, we illustrate how subtle differences\nin the underlying game formulations of existing methods can cause large\ndifferences in the attributions for a prediction. We then present a general\ngame formulation that unifies existing methods, and enables straightforward\nconfidence intervals on their attributions. Furthermore, it allows us to\ninterpret the attributions as contrastive explanations of an input relative to\na distribution of reference inputs. We tie this idea to classic research in\ncognitive psychology on contrastive explanations, and propose a conceptual\nframework for generating and interpreting explanations for ML models, called\nformulate, approximate, explain (FAE). We apply this framework to explain\nblack-box models trained on two UCI datasets and a Lending Club dataset.\n
Paper
References (27)
Scroll for more · 15 remaining