Machine translation (MT) has benefited from using synthetic training data\noriginating from translating monolingual corpora, a technique known as\nbacktranslation. Combining backtranslated data from different sources has led\nto better results than when using such data in isolation. In this work we\nanalyse the impact that data translated with rule-based, phrase-based\nstatistical and neural MT systems has on new MT systems. We use a real-world\nlow-resource use-case (Basque-to-Spanish in the clinical domain) as well as a\nhigh-resource language pair (German-to-English) to test different scenarios\nwith backtranslation and employ data selection to optimise the synthetic\ncorpora. We exploit different data selection strategies in order to reduce the\namount of data used, while at the same time maintaining high-quality MT\nsystems. We further tune the data selection method by taking into account the\nquality of the MT systems used for backtranslation and lexical diversity of the\nresulting corpora. Our experiments show that incorporating backtranslated data\nfrom different sources can be beneficial, and that availing of data selection\ncan yield improved performance.\n
Paper
References (50)
Scroll for more · 38 remaining