Unsupervised translation has reached impressive performance on resource-rich\nlanguage pairs such as English-French and English-German. However, early\nstudies have shown that in more realistic settings involving low-resource, rare\nlanguages, unsupervised translation performs poorly, achieving less than 3.0\nBLEU. In this work, we show that multilinguality is critical to making\nunsupervised systems practical for low-resource settings. In particular, we\npresent a single model for 5 low-resource languages (Gujarati, Kazakh, Nepali,\nSinhala, and Turkish) to and from English directions, which leverages\nmonolingual and auxiliary parallel data from other high-resource language pairs\nvia a three-stage training scheme. We outperform all current state-of-the-art\nunsupervised baselines for these languages, achieving gains of up to 14.4 BLEU.\nAdditionally, we outperform a large collection of supervised WMT submissions\nfor various language pairs as well as match the performance of the current\nstate-of-the-art supervised model for Nepali-English. We conduct a series of\nablation studies to establish the robustness of our model under different\ndegrees of data quality, as well as to analyze the factors which led to the\nsuperior performance of the proposed approach over traditional unsupervised\nmodels.\n
Paper
References (53)
Scroll for more · 38 remaining