BERT is Robust! A Case Against Synonym-Based Adversarial Examples in Text Classification

Deep Neural Networks have taken Natural Language Processing by storm. While\nthis led to incredible improvements across many tasks, it also initiated a new\nresearch field, questioning the robustness of these neural networks by\nattacking them. In this paper, we investigate four word substitution-based\nattacks on BERT. We combine a human evaluation of individual word substitutions\nand a probabilistic analysis to show that between 96% and 99% of the analyzed\nattacks do not preserve semantics, indicating that their success is mainly\nbased on feeding poor data to the model. To further confirm that, we introduce\nan efficient data augmentation procedure and show that many adversarial\nexamples can be prevented by including data similar to the attacks during\ntraining. An additional post-processing step reduces the success rates of\nstate-of-the-art attacks below 5%. Finally, by looking at more reasonable\nthresholds on constraints for word substitutions, we conclude that BERT is a\nlot more robust than research on attacks suggests.\n

Paper

References (38)

Scroll for more · 26 remaining

Similar papers

© 2026 NYSGPT2525 LLC