Is MAP Decoding All You Need? The Inadequacy of the Mode in Neural Machine Translation

Recent studies have revealed a number of pathologies of neural machine\ntranslation (NMT) systems. Hypotheses explaining these mostly suggest there is\nsomething fundamentally wrong with NMT as a model or its training algorithm,\nmaximum likelihood estimation (MLE). Most of this evidence was gathered using\nmaximum a posteriori (MAP) decoding, a decision rule aimed at identifying the\nhighest-scoring translation, i.e. the mode. We argue that the evidence\ncorroborates the inadequacy of MAP decoding more than casts doubt on the model\nand its training algorithm. In this work, we show that translation\ndistributions do reproduce various statistics of the data well, but that beam\nsearch strays from such statistics. We show that some of the known pathologies\nand biases of NMT are due to MAP decoding and not to NMT's statistical\nassumptions nor MLE. In particular, we show that the most likely translations\nunder the model accumulate so little probability mass that the mode can be\nconsidered essentially arbitrary. We therefore advocate for the use of decision\nrules that take into account the translation distribution holistically. We show\nthat an approximation to minimum Bayes risk decoding gives competitive results\nconfirming that NMT models do capture important aspects of translation well in\nexpectation.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC