Measuring and Increasing Context Usage in Context-Aware Machine Translation

Recent work in neural machine translation has demonstrated both the necessity\nand feasibility of using inter-sentential context -- context from sentences\nother than those currently being translated. However, while many current\nmethods present model architectures that theoretically can use this extra\ncontext, it is often not clear how much they do actually utilize it at\ntranslation time. In this paper, we introduce a new metric, conditional\ncross-mutual information, to quantify the usage of context by these models.\nUsing this metric, we measure how much document-level machine translation\nsystems use particular varieties of context. We find that target context is\nreferenced more than source context, and that conditioning on a longer context\nhas a diminishing effect on results. We then introduce a new, simple training\nmethod, context-aware word dropout, to increase the usage of context by\ncontext-aware models. Experiments show that our method increases context usage\nand that this reflects on the translation quality according to metrics such as\nBLEU and COMET, as well as performance on anaphoric pronoun resolution and\nlexical cohesion contrastive datasets.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC