Distill, Adapt, Distill: Training Small, In-Domain Models for Neural Machine Translation

We explore best practices for training small, memory efficient machine\ntranslation models with sequence-level knowledge distillation in the domain\nadaptation setting. While both domain adaptation and knowledge distillation are\nwidely-used, their interaction remains little understood. Our large-scale\nempirical results in machine translation (on three language pairs with three\ndomains each) suggest distilling twice for best performance: once using\ngeneral-domain data and again using in-domain data with an adapted teacher.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC