Fine-Grained Neural Network Explanation by Identifying Input Features with Predictive Information

One principal approach for illuminating a black-box neural network is feature\nattribution, i.e. identifying the importance of input features for the\nnetwork's prediction. The predictive information of features is recently\nproposed as a proxy for the measure of their importance. So far, the predictive\ninformation is only identified for latent features by placing an information\nbottleneck within the network. We propose a method to identify features with\npredictive information in the input domain. The method results in fine-grained\nidentification of input features' information and is agnostic to network\narchitecture. The core idea of our method is leveraging a bottleneck on the\ninput that only lets input features associated with predictive latent features\npass through. We compare our method with several feature attribution methods\nusing mainstream feature attribution evaluation experiments. The code is\npublicly available.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC