An Empirical Study of Derivative-Free-Optimization Algorithms for Targeted Black-Box Attacks in Deep Neural Networks
We perform a comprehensive study on the performance of derivative free\noptimization (DFO) algorithms for the generation of targeted black-box\nadversarial attacks on Deep Neural Network (DNN) classifiers assuming the\nperturbation energy is bounded by an $\\ell_\\infty$ constraint and the number of\nqueries to the network is limited. This paper considers four pre-existing\nstate-of-the-art DFO-based algorithms along with the introduction of a new\nalgorithm built on BOBYQA, a model-based DFO method. We compare these\nalgorithms in a variety of settings according to the fraction of images that\nthey successfully misclassify given a maximum number of queries to the DNN.\n The experiments disclose how the likelihood of finding an adversarial example\ndepends on both the algorithm used and the setting of the attack; algorithms\nlimiting the search of adversarial example to the vertices of the $\\ell^\\infty$\nconstraint work particularly well without structural defenses, while the\npresented BOBYQA based algorithm works better for especially small perturbation\nenergies. This variance in performance highlights the importance of new\nalgorithms being compared to the state-of-the-art in a variety of settings, and\nthe effectiveness of adversarial defenses being tested using as wide a range of\nalgorithms as possible.\n
Paper
References (51)
Scroll for more · 38 remaining