Contrasting Human- and Machine-Generated Word-Level Adversarial Examples for Text Classification
Research shows that natural language processing models are generally\nconsidered to be vulnerable to adversarial attacks; but recent work has drawn\nattention to the issue of validating these adversarial inputs against certain\ncriteria (e.g., the preservation of semantics and grammaticality). Enforcing\nconstraints to uphold such criteria may render attacks unsuccessful, raising\nthe question of whether valid attacks are actually feasible. In this work, we\ninvestigate this through the lens of human language ability. We report on\ncrowdsourcing studies in which we task humans with iteratively modifying words\nin an input text, while receiving immediate model feedback, with the aim of\ncausing a sentiment classification model to misclassify the example. Our\nfindings suggest that humans are capable of generating a substantial amount of\nadversarial examples using semantics-preserving word substitutions. We analyze\nhow human-generated adversarial examples compare to the recently proposed\nTextFooler, Genetic, BAE and SememePSO attack algorithms on the dimensions\nnaturalness, preservation of sentiment, grammaticality and substitution rate.\nOur findings suggest that human-generated adversarial examples are not more\nable than the best algorithms to generate natural-reading, sentiment-preserving\nexamples, though they do so by being much more computationally efficient.\n
Paper
References (35)
Scroll for more · 23 remaining