Generalised Unsupervised Domain Adaptation of Neural Machine Translation with Cross-Lingual Data Selection

This paper considers the unsupervised domain adaptation problem for neural\nmachine translation (NMT), where we assume the access to only monolingual text\nin either the source or target language in the new domain. We propose a\ncross-lingual data selection method to extract in-domain sentences in the\nmissing language side from a large generic monolingual corpus. Our proposed\nmethod trains an adaptive layer on top of multilingual BERT by contrastive\nlearning to align the representation between the source and target language.\nThis then enables the transferability of the domain classifier between the\nlanguages in a zero-shot manner. Once the in-domain data is detected by the\nclassifier, the NMT model is then adapted to the new domain by jointly learning\ntranslation and domain discrimination tasks. We evaluate our cross-lingual data\nselection method on NMT across five diverse domains in three language pairs, as\nwell as a real-world scenario of translation for COVID-19. The results show\nthat our proposed method outperforms other selection baselines up to +1.5 BLEU\nscore.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC