Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial Training

Delusive attacks aim to substantially deteriorate the test accuracy of the\nlearning model by slightly perturbing the features of correctly labeled\ntraining examples. By formalizing this malicious attack as finding the\nworst-case training data within a specific $\\infty$-Wasserstein ball, we show\nthat minimizing adversarial risk on the perturbed data is equivalent to\noptimizing an upper bound of natural risk on the original data. This implies\nthat adversarial training can serve as a principled defense against delusive\nattacks. Thus, the test accuracy decreased by delusive attacks can be largely\nrecovered by adversarial training. To further understand the internal mechanism\nof the defense, we disclose that adversarial training can resist the delusive\nperturbations by preventing the learner from overly relying on non-robust\nfeatures in a natural setting. Finally, we complement our theoretical findings\nwith a set of experiments on popular benchmark datasets, which show that the\ndefense withstands six different practical attacks. Both theoretical and\nempirical results vote for adversarial training when confronted with delusive\nadversaries.\n

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC