When Feature Attribution Methods Meet Robust Models: A Survey and Outlook

The black-box nature of deep neural networks has motivated the development of explanation methods to improve model transparency. Among them, feature attribution (FA) is widely adopted for its intuitive outputs and broad applicability. While FA methods have been extensively evaluated on standard models, their behavior on adversarially robust models remains unclear. To address this, we present the first systematic study of FA methods under robust training, focusing on five key properties: fidelity, stability, conciseness, class sensitivity, and weight sensitivity. Through quantitative evaluation of six FA methods across four training strategies, i.e., standard training, adversarial training (AT), adversarial robustness distillation (ARD), and adversarial robustness regularization (ARR), we reveal several key findings: (1) robustness significantly improves attribution fidelity and stability; (2) white-box methods yield more concise explanations under AT and ARD; and (3) robust models reduce class sensitivity and weight sensitivity, indicating a trade-off between robustness and interpretability. These insights highlight both the benefits and limitations of applying FA to robust models. We conclude by outlining directions for robustness-aware attribution techniques and joint training strategies that balance interpretability and security, paving the way toward trustworthy and explainable artificial intelligence systems.<br/>

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC