Building a Multi-domain Neural Machine Translation Model using Knowledge Distillation

Lack of specialized data makes building a multi-domain neural machine\ntranslation tool challenging. Although emerging literature dealing with low\nresource languages starts to show promising results, most state-of-the-art\nmodels used millions of sentences. Today, the majority of multi-domain\nadaptation techniques are based on complex and sophisticated architectures that\nare not adapted for real-world applications. So far, no scalable method is\nperforming better than the simple yet effective mixed-finetuning, i.e\nfinetuning a generic model with a mix of all specialized data and generic data.\nIn this paper, we propose a new training pipeline where knowledge distillation\nand multiple specialized teachers allow us to efficiently finetune a model\nwithout adding new costs at inference time. Our experiments demonstrated that\nour training pipeline allows improving the performance of multi-domain\ntranslation over finetuning in configurations with 2, 3, and 4 domains by up to\n2 points in BLEU.\n

Paper

References (28)

Scroll for more · 16 remaining

Similar papers

© 2026 NYSGPT2525 LLC