Explaining Explanations: Axiomatic Feature Interactions for Deep Networks

Recent work has shown great promise in explaining neural network behavior. In\nparticular, feature attribution methods explain which features were most\nimportant to a model's prediction on a given input. However, for many tasks,\nsimply knowing which features were important to a model's prediction may not\nprovide enough insight to understand model behavior. The interactions between\nfeatures within the model may better help us understand not only the model, but\nalso why certain features are more important than others. In this work, we\npresent Integrated Hessians, an extension of Integrated Gradients that explains\npairwise feature interactions in neural networks. Integrated Hessians overcomes\nseveral theoretical limitations of previous methods to explain interactions,\nand unlike such previous methods is not limited to a specific architecture or\nclass of neural network. Additionally, we find that our method is faster than\nexisting methods when the number of features is large, and outperforms previous\nmethods on existing quantitative benchmarks. Code available at\nhttps://github.com/suinleelab/path_explain\n

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC