Despite our best efforts, deep learning models remain highly vulnerable to\neven tiny adversarial perturbations applied to the inputs. The ability to\nextract information from solely the output of a machine learning model to craft\nadversarial perturbations to black-box models is a practical threat against\nreal-world systems, such as autonomous cars or machine learning models exposed\nas a service (MLaaS). Of particular interest are sparse attacks. The\nrealization of sparse attacks in black-box models demonstrates that machine\nlearning models are more vulnerable than we believe. Because these attacks aim\nto minimize the number of perturbed pixels measured by l_0 norm-required to\nmislead a model by solely observing the decision (the predicted label) returned\nto a model query; the so-called decision-based attack setting. But, such an\nattack leads to an NP-hard optimization problem. We develop an evolution-based\nalgorithm-SparseEvo-for the problem and evaluate against both convolutional\ndeep neural networks and vision transformers. Notably, vision transformers are\nyet to be investigated under a decision-based attack setting. SparseEvo\nrequires significantly fewer model queries than the state-of-the-art sparse\nattack Pointwise for both untargeted and targeted attacks. The attack\nalgorithm, although conceptually simple, is also competitive with only a\nlimited query budget against the state-of-the-art gradient-based whitebox\nattacks in standard computer vision tasks such as ImageNet. Importantly, the\nquery efficient SparseEvo, along with decision-based attacks, in general, raise\nnew questions regarding the safety of deployed systems and poses new directions\nto study and understand the robustness of machine learning models.\n
Paper
References (42)
Scroll for more · 30 remaining