Leveraging Discourse Rewards for Document-Level Neural Machine Translation

Document-level machine translation focuses on the translation of entire\ndocuments from a source to a target language. It is widely regarded as a\nchallenging task since the translation of the individual sentences in the\ndocument needs to retain aspects of the discourse at document level. However,\ndocument-level translation models are usually not trained to explicitly ensure\ndiscourse quality. Therefore, in this paper we propose a training approach that\nexplicitly optimizes two established discourse metrics, lexical cohesion (LC)\nand coherence (COH), by using a reinforcement learning objective. Experiments\nover four different language pairs and three translation domains have shown\nthat our training approach has been able to achieve more cohesive and coherent\ndocument translations than other competitive approaches, yet without\ncompromising the faithfulness to the reference translation. In the case of the\nZh-En language pair, our method has achieved an improvement of 2.46 percentage\npoints (pp) in LC and 1.17 pp in COH over the runner-up, while at the same time\nimproving 0.63 pp in BLEU score and 0.47 pp in F_BERT.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC