BERT-Defense: A Probabilistic Model Based on BERT to Combat Cognitively Inspired Orthographic Adversarial Attacks
Adversarial attacks expose important blind spots of deep learning systems.\nWhile word- and sentence-level attack scenarios mostly deal with finding\nsemantic paraphrases of the input that fool NLP models, character-level attacks\ntypically insert typos into the input stream. It is commonly thought that these\nare easier to defend via spelling correction modules. In this work, we show\nthat both a standard spellchecker and the approach of Pruthi et al. (2019),\nwhich trains to defend against insertions, deletions and swaps, perform poorly\non the character-level benchmark recently proposed in Eger and Benz (2020)\nwhich includes more challenging attacks such as visual and phonetic\nperturbations and missing word segmentations. In contrast, we show that an\nuntrained iterative approach which combines context-independent character-level\ninformation with context-dependent information from BERT's masked language\nmodeling can perform on par with human crowd-workers from Amazon Mechanical\nTurk (AMT) supervised via 3-shot learning.\n
Paper
References (36)
Scroll for more · 24 remaining