Generating End-to-End Adversarial Examples for Malware Classifiers Using Explainability

In recent years, the topic of explainable machine learning (ML) has been\nextensively researched. Up until now, this research focused on regular ML users\nuse-cases such as debugging a ML model. This paper takes a different posture\nand show that adversaries can leverage explainable ML to bypass multi-feature\ntypes malware classifiers. Previous adversarial attacks against such\nclassifiers only add new features and not modify existing ones to avoid harming\nthe modified malware executable's functionality. Current attacks use a single\nalgorithm that both selects which features to modify and modifies them blindly,\ntreating all features the same. In this paper, we present a different approach.\nWe split the adversarial example generation task into two parts: First we find\nthe importance of all features for a specific sample using explainability\nalgorithms, and then we conduct a feature-specific modification,\nfeature-by-feature. In order to apply our attack in black-box scenarios, we\nintroduce the concept of transferability of explainability, that is, applying\nexplainability algorithms to different classifiers using different features\nsubsets and trained on different datasets still result in a similar subset of\nimportant features. We conclude that explainability algorithms can be leveraged\nby adversaries and thus the advocates of training more interpretable\nclassifiers should consider the trade-off of higher vulnerability of those\nclassifiers to adversarial attacks.\n

Paper

References (39)

Scroll for more · 27 remaining

Similar papers

© 2026 NYSGPT2525 LLC