There has been significant progress in recent years in the field of Natural\nLanguage Processing thanks to the introduction of the Transformer architecture.\nCurrent state-of-the-art models, via a large number of parameters and\npre-training on massive text corpus, have shown impressive results on several\ndownstream tasks. Many researchers have studied previous (non-Transformer)\nmodels to understand their actual behavior under different scenarios, showing\nthat these models are taking advantage of clues or failures of datasets and\nthat slight perturbations on the input data can severely reduce their\nperformance. In contrast, recent models have not been systematically tested\nwith adversarial-examples in order to show their robustness under severe stress\nconditions. For that reason, this work evaluates three Transformer-based models\n(RoBERTa, XLNet, and BERT) in Natural Language Inference (NLI) and Question\nAnswering (QA) tasks to know if they are more robust or if they have the same\nflaws as their predecessors. As a result, our experiments reveal that RoBERTa,\nXLNet and BERT are more robust than recurrent neural network models to stress\ntests for both NLI and QA tasks. Nevertheless, they are still very fragile and\ndemonstrate various unexpected behaviors, thus revealing that there is still\nroom for future improvement in this field.\n
Paper
References (41)
Scroll for more · 29 remaining