Towards Understanding the Regularization of Adversarial Robustness on Neural Networks

The problem of adversarial examples has shown that modern Neural Network (NN)\nmodels could be rather fragile. Among the more established techniques to solve\nthe problem, one is to require the model to be {\\it $\\epsilon$-adversarially\nrobust} (AR); that is, to require the model not to change predicted labels when\nany given input examples are perturbed within a certain range. However, it is\nobserved that such methods would lead to standard performance degradation,\ni.e., the degradation on natural examples. In this work, we study the\ndegradation through the regularization perspective. We identify quantities from\ngeneralization analysis of NNs; with the identified quantities we empirically\nfind that AR is achieved by regularizing/biasing NNs towards less confident\nsolutions by making the changes in the feature space (induced by changes in the\ninstance space) of most layers smoother uniformly in all directions; so to a\ncertain extent, it prevents sudden change in prediction w.r.t. perturbations.\nHowever, the end result of such smoothing concentrates samples around decision\nboundaries, resulting in less confident solutions, and leads to worse standard\nperformance. Our studies suggest that one might consider ways that build AR\ninto NNs in a gentler way to avoid the problematic regularization.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC