We describe an unsupervised method to create pseudo-parallel corpora for\nmachine translation (MT) from unaligned text. We use multilingual BERT to\ncreate source and target sentence embeddings for nearest-neighbor search and\nadapt the model via self-training. We validate our technique by extracting\nparallel sentence pairs on the BUCC 2017 bitext mining task and observe up to a\n24.5 point increase (absolute) in F1 scores over previous unsupervised methods.\nWe then improve an XLM-based unsupervised neural MT system pre-trained on\nWikipedia by supplementing it with pseudo-parallel text mined from the same\ncorpus, boosting unsupervised translation performance by up to 3.5 BLEU on the\nWMT'14 French-English and WMT'16 German-English tasks and outperforming the\nprevious state-of-the-art. Finally, we enrich the IWSLT'15 English-Vietnamese\ncorpus with pseudo-parallel Wikipedia sentence pairs, yielding a 1.2 BLEU\nimprovement on the low-resource MT task. We demonstrate that unsupervised\nbitext mining is an effective way of augmenting MT datasets and complements\nexisting techniques like initializing with pre-trained contextual embeddings.\n
Paper
References (50)
Scroll for more · 38 remaining