In recent years, the topic of explainable machine learning (ML) has been\nextensively researched. Up until now, this research focused on regular ML users\nuse-cases such as debugging a ML model. This paper takes a different posture\nand show that adversaries can leverage explainable ML to bypass multi-feature\ntypes malware classifiers. Previous adversarial attacks against such\nclassifiers only add new features and not modify existing ones to avoid harming\nthe modified malware executable's functionality. Current attacks use a single\nalgorithm that both selects which features to modify and modifies them blindly,\ntreating all features the same. In this paper, we present a different approach.\nWe split the adversarial example generation task into two parts: First we find\nthe importance of all features for a specific sample using explainability\nalgorithms, and then we conduct a feature-specific modification,\nfeature-by-feature. In order to apply our attack in black-box scenarios, we\nintroduce the concept of transferability of explainability, that is, applying\nexplainability algorithms to different classifiers using different features\nsubsets and trained on different datasets still result in a similar subset of\nimportant features. We conclude that explainability algorithms can be leveraged\nby adversaries and thus the advocates of training more interpretable\nclassifiers should consider the trade-off of higher vulnerability of those\nclassifiers to adversarial attacks.\n
Paper
References (39)
Scroll for more · 27 remaining