Adversarial perturbations dramatically decrease the accuracy of\nstate-of-the-art image classifiers. In this paper, we propose and analyze a\nsimple and computationally efficient defense strategy: inject random Gaussian\nnoise, discretize each pixel, and then feed the result into any pre-trained\nclassifier. Theoretically, we show that our randomized discretization strategy\nreduces the KL divergence between original and adversarial inputs, leading to a\nlower bound on the classification accuracy of any classifier against any\n(potentially whitebox) $\\ell_\\infty$-bounded adversarial attack. Empirically,\nwe evaluate our defense on adversarial examples generated by a strong iterative\nPGD attack. On ImageNet, our defense is more robust than adversarially-trained\nnetworks and the winning defenses of the NIPS 2017 Adversarial Attacks &\nDefenses competition.\n
Paper
References (42)
Scroll for more · 30 remaining