High-performance neural language models have obtained state-of-the-art\nresults on a wide range of Natural Language Processing (NLP) tasks. However,\nresults for common benchmark datasets often do not reflect model reliability\nand robustness when applied to noisy, real-world data. In this study, we design\nand implement various types of character-level and word-level perturbation\nmethods to simulate realistic scenarios in which input texts may be slightly\nnoisy or different from the data distribution on which NLP systems were\ntrained. Conducting comprehensive experiments on different NLP tasks, we\ninvestigate the ability of high-performance language models such as BERT,\nXLNet, RoBERTa, and ELMo in handling different types of input perturbations.\nThe results suggest that language models are sensitive to input perturbations\nand their performance can decrease even when small changes are introduced. We\nhighlight that models need to be further improved and that current benchmarks\nare not reflecting model robustness well. We argue that evaluations on\nperturbed inputs should routinely complement widely-used benchmarks in order to\nyield a more realistic understanding of NLP systems robustness.\n
Paper
References (37)
Scroll for more · 25 remaining