Using Shapley Values and Variational Autoencoders to Explain Predictive Models with Dependent Mixed Features
Shapley values are today extensively used as a model-agnostic explanation\nframework to explain complex predictive machine learning models. Shapley values\nhave desirable theoretical properties and a sound mathematical foundation in\nthe field of cooperative game theory. Precise Shapley value estimates for\ndependent data rely on accurate modeling of the dependencies between all\nfeature combinations. In this paper, we use a variational autoencoder with\narbitrary conditioning (VAEAC) to model all feature dependencies\nsimultaneously. We demonstrate through comprehensive simulation studies that\nour VAEAC approach to Shapley value estimation outperforms the state-of-the-art\nmethods for a wide range of settings for both continuous and mixed dependent\nfeatures. For high-dimensional settings, our VAEAC approach with a non-uniform\nmasking scheme significantly outperforms competing methods. Finally, we apply\nour VAEAC approach to estimate Shapley value explanations for the Abalone data\nset from the UCI Machine Learning Repository.\n