Despite the impressive performance of neural language models in natural language processing, they remain vulnerable to adversarial attacks, particularly those based on synonym substitutions at the word level. Recent defense methods, however, often generate worst-case adversarial examples or upper bounds on propagation by independently perturbing each data sample. This approach limits the model’s ability to generalize to unseen data. To address this limitation, we propose a distributional robustness framework to defend against synonym-based word-level adversarial attacks. Our framework identifies the worst-case distribution within a known uncertainty set to craft adversarial examples and models the solution space as the convex hull of word vectors. This convex hull is sufficiently inclusive to cover all potential substitutions while excluding unnecessary ones. Furthermore, we provide an efficiently computable, data-dependent upper bound on the worst-case loss, ensuring that the worst-case performance of the output from our principled adversarial training procedure does not exceed this bound. Additionally, we certify the generalizability of the robustness level to the training set. Experiments on benchmark datasets demonstrate that our framework achieves comparable adversarial robustness to other defense methods under synonym-based word-level attacks.
Paper
Full text
Distributional Robustness Framework against Word-level Adversarial Attacks
Semantic Scholar · 2025
Abstract
Despite the impressive performance of neural language models in natural language processing, they remain vulnerable to adversarial attacks, particularly those based on synonym substitutions at the word level. Recent defense methods, however, often generate worst-case adversarial examples or upper bounds on propagation by independently perturbing each data sample. This approach limits the model’s ability to generalize to unseen data. To address this limitation, we propose a distributional robustness framework to defend against synonym-based word-level adversarial attacks. Our framework identifies the worst-case distribution within a known uncertainty set to craft adversarial examples and models the solution space as the convex hull of word vectors. This convex hull is sufficiently inclusive to cover all potential substitutions while excluding unnecessary ones. Furthermore, we provide an efficiently computable, data-dependent upper bound on the worst-case loss, ensuring that the worst-case performance of the output from our principled adversarial training procedure does not exceed this bound. Additionally, we certify the generalizability of the robustness level to the training set. Experiments on benchmark datasets demonstrate that our framework achieves comparable adversarial robustness to other defense methods under synonym-based word-level attacks.