Recent efforts have shown that neural text processing models are vulnerable\nto adversarial examples, but the nature of these examples is poorly understood.\nIn this work, we show that adversarial attacks against CNN, LSTM and\nTransformer-based classification models perform word substitutions that are\nidentifiable through frequency differences between replaced words and their\ncorresponding substitutions. Based on these findings, we propose\nfrequency-guided word substitutions (FGWS), a simple algorithm exploiting the\nfrequency properties of adversarial word substitutions for the detection of\nadversarial examples. FGWS achieves strong performance by accurately detecting\nadversarial examples on the SST-2 and IMDb sentiment datasets, with F1\ndetection scores of up to 91.4% against RoBERTa-based classification models. We\ncompare our approach against a recently proposed perturbation discrimination\nframework and show that we outperform it by up to 13.0% F1.\n
Paper
References (61)
Scroll for more · 38 remaining