The growing incorporation of artificial neural networks (NNs) into many\nfields, and especially into life-critical systems, is restrained by their\nvulnerability to adversarial examples (AEs). Some existing defense methods can\nincrease NNs' robustness, but they often require special architecture or\ntraining procedures and are irrelevant to already trained models. In this\npaper, we propose a simple defense that combines feature visualization with\ninput modification, and can, therefore, be applicable to various pre-trained\nnetworks. By reviewing several interpretability methods, we gain new insights\nregarding the influence of AEs on NNs' computation. Based on that, we\nhypothesize that information about the "true" object is preserved within the\nNN's activity, even when the input is adversarial, and present a feature\nvisualization version that can extract that information in the form of\nrelevance heatmaps. We then use these heatmaps as a basis for our defense, in\nwhich the adversarial effects are corrupted by massive blurring. We also\nprovide a new evaluation metric that can capture the effects of both attacks\nand defenses more thoroughly and descriptively, and demonstrate the\neffectiveness of the defense and the utility of the suggested evaluation\nmeasurement with VGG19 results on the ImageNet dataset.\n
Paper
References (33)
Scroll for more · 21 remaining