Do Gradient-based Explanations Tell Anything About Adversarial Robustness to Android Malware?
While machine-learning algorithms have demonstrated a strong ability in\ndetecting Android malware, they can be evaded by sparse evasion attacks crafted\nby injecting a small set of fake components, e.g., permissions and system\ncalls, without compromising intrusive functionality. Previous work has shown\nthat, to improve robustness against such attacks, learning algorithms should\navoid overemphasizing few discriminant features, providing instead decisions\nthat rely upon a large subset of components. In this work, we investigate\nwhether gradient-based attribution methods, used to explain classifiers'\ndecisions by identifying the most relevant features, can be used to help\nidentify and select more robust algorithms. To this end, we propose to exploit\ntwo different metrics that represent the evenness of explanations, and a new\ncompact security measure called Adversarial Robustness Metric. Our experiments\nconducted on two different datasets and five classification algorithms for\nAndroid malware detection show that a strong connection exists between the\nuniformity of explanations and adversarial robustness. In particular, we found\nthat popular techniques like Gradient*Input and Integrated Gradients are\nstrongly correlated to security when applied to both linear and nonlinear\ndetectors, while more elementary explanation techniques like the simple\nGradient do not provide reliable information about the robustness of such\nclassifiers.\n
Paper
References (73)
Scroll for more · 38 remaining