Adversarial attacks based on randomized search schemes have obtained\nstate-of-the-art results in black-box robustness evaluation recently. However,\nas we demonstrate in this work, their efficiency in different query budget\nregimes depends on manual design and heuristic tuning of the underlying\nproposal distributions. We study how this issue can be addressed by adapting\nthe proposal distribution online based on the information obtained during the\nattack. We consider Square Attack, which is a state-of-the-art score-based\nblack-box attack, and demonstrate how its performance can be improved by a\nlearned controller that adjusts the parameters of the proposal distribution\nonline during the attack. We train the controller using gradient-based\nend-to-end training on a CIFAR10 model with white box access. We demonstrate\nthat plugging the learned controller into the attack consistently improves its\nblack-box robustness estimate in different query regimes by up to 20% for a\nwide range of different models with black-box access. We further show that the\nlearned adaptation principle transfers well to the other data distributions\nsuch as CIFAR100 or ImageNet and to the targeted attack setting.\n
Paper
References (78)
Scroll for more · 38 remaining