Breaking certified defenses: Semantic adversarial examples with spoofed robustness certificates

To deflect adversarial attacks, a range of "certified" classifiers have been\nproposed. In addition to labeling an image, certified classifiers produce (when\npossible) a certificate guaranteeing that the input image is not an\n$\\ell_p$-bounded adversarial example. We present a new attack that exploits not\nonly the labelling function of a classifier, but also the certificate\ngenerator. The proposed method applies large perturbations that place images\nfar from a class boundary while maintaining the imperceptibility property of\nadversarial examples. The proposed "Shadow Attack" causes certifiably robust\nnetworks to mislabel an image and simultaneously produce a "spoofed"\ncertificate of robustness.\n

Paper

References (36)

Scroll for more · 24 remaining

Similar papers

© 2026 NYSGPT2525 LLC