Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing

We frame the task of machine translation evaluation as one of scoring machine\ntranslation output with a sequence-to-sequence paraphraser, conditioned on a\nhuman reference. We propose training the paraphraser as a multilingual NMT\nsystem, treating paraphrasing as a zero-shot translation task (e.g., Czech to\nCzech). This results in the paraphraser's output mode being centered around a\ncopy of the input sequence, which represents the best case scenario where the\nMT system output matches a human reference. Our method is simple and intuitive,\nand does not require human judgements for training. Our single model (trained\nin 39 languages) outperforms or statistically ties with all prior metrics on\nthe WMT 2019 segment-level shared metrics task in all languages (excluding\nGujarati where the model had no training data). We also explore using our model\nfor the task of quality estimation as a metric--conditioning on the source\ninstead of the reference--and find that it significantly outperforms every\nsubmission to the WMT 2019 shared task on quality estimation in every language\npair.\n

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC