BET: A Backtranslation Approach for Easy Data Augmentation in Transformer-based Paraphrase Identification Context
Newly-introduced deep learning architectures, namely BERT, XLNet, RoBERTa and\nALBERT, have been proved to be robust on several NLP tasks. However, the\ndatasets trained on these architectures are fixed in terms of size and\ngeneralizability. To relieve this issue, we apply one of the most inexpensive\nsolutions to update these datasets. We call this approach BET by which we\nanalyze the backtranslation data augmentation on the transformer-based\narchitectures. Using the Google Translate API with ten intermediary languages\nfrom ten different language families, we externally evaluate the results in the\ncontext of automatic paraphrase identification in a transformer-based\nframework. Our findings suggest that BET improves the paraphrase identification\nperformance on the Microsoft Research Paraphrase Corpus (MRPC) to more than 3%\non both accuracy and F1 score. We also analyze the augmentation in the low-data\nregime with downsampled versions of MRPC, Twitter Paraphrase Corpus (TPC) and\nQuora Question Pairs. In many low-data cases, we observe a switch from a\nfailing model on the test set to reasonable performances. The results\ndemonstrate that BET is a highly promising data augmentation technique: to push\nthe current state-of-the-art of existing datasets and to bootstrap the\nutilization of deep learning architectures in the low-data regime of a hundred\nsamples.\n
Paper
References (40)
Scroll for more · 28 remaining