Metrics and methods for robustness evaluation of neural networks with generative models

Recent studies have shown that modern deep neural network classifiers are\neasy to fool, assuming that an adversary is able to slightly modify their\ninputs. Many papers have proposed adversarial attacks, defenses and methods to\nmeasure robustness to such adversarial perturbations. However, most commonly\nconsidered adversarial examples are based on $\\ell_p$-bounded perturbations in\nthe input space of the neural network, which are unlikely to arise naturally.\nRecently, especially in computer vision, researchers discovered "natural" or\n"semantic" perturbations, such as rotations, changes of brightness, or more\nhigh-level changes, but these perturbations have not yet been systematically\nutilized to measure the performance of classifiers. In this paper, we propose\nseveral metrics to measure robustness of classifiers to natural adversarial\nexamples, and methods to evaluate them. These metrics, called latent space\nperformance metrics, are based on the ability of generative models to capture\nprobability distributions, and are defined in their latent spaces. On three\nimage classification case studies, we evaluate the proposed metrics for several\nclassifiers, including ones trained in conventional and robust ways. We find\nthat the latent counterparts of adversarial robustness are associated with the\naccuracy of the classifier rather than its conventional adversarial robustness,\nbut the latter is still reflected on the properties of found latent\nperturbations. In addition, our novel method of finding latent adversarial\nperturbations demonstrates that these perturbations are often perceptually\nsmall.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC