Interpretable Neural Networks with Frank-Wolfe: Sparse Relevance Maps and Relevance Orderings

We study the effects of constrained optimization formulations and Frank-Wolfe\nalgorithms for obtaining interpretable neural network predictions.\nReformulating the Rate-Distortion Explanations (RDE) method for relevance\nattribution as a constrained optimization problem provides precise control over\nthe sparsity of relevance maps. This enables a novel multi-rate as well as a\nrelevance-ordering variant of RDE that both empirically outperform standard RDE\nand other baseline methods in a well-established comparison test. We showcase\nseveral deterministic and stochastic variants of the Frank-Wolfe algorithm and\ntheir effectiveness for RDE.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC