Improving performance of deep learning models with axiomatic attribution priors and expected gradients

Recent research has demonstrated that feature attribution methods for deep\nnetworks can themselves be incorporated into training; these attribution priors\noptimize for a model whose attributions have certain desirable properties --\nmost frequently, that particular features are important or unimportant. These\nattribution priors are often based on attribution methods that are not\nguaranteed to satisfy desirable interpretability axioms, such as completeness\nand implementation invariance. Here, we introduce attribution priors to\noptimize for higher-level properties of explanations, such as smoothness and\nsparsity, enabled by a fast new attribution method formulation called expected\ngradients that satisfies many important interpretability axioms. This improves\nmodel performance on many real-world tasks where previous attribution priors\nfail. Our experiments show that the gains from combining higher-level\nattribution priors with expected gradients attributions are consistent across\nimage, gene expression, and health care data sets. We believe this work\nmotivates and provides the necessary tools to support the widespread adoption\nof axiomatic attribution priors in many areas of applied machine learning. The\nimplementations and our results have been made freely available to academic\ncommunities.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC