Online Learning Meets Machine Translation Evaluation: Finding the Best Systems with the Least Human Effort
In Machine Translation, assessing the quality of a large amount of automatic\ntranslations can be challenging. Automatic metrics are not reliable when it\ncomes to high performing systems. In addition, resorting to human evaluators\ncan be expensive, especially when evaluating multiple systems. To overcome the\nlatter challenge, we propose a novel application of online learning that, given\nan ensemble of Machine Translation systems, dynamically converges to the best\nsystems, by taking advantage of the human feedback available. Our experiments\non WMT'19 datasets show that our online approach quickly converges to the top-3\nranked systems for the language pairs considered, despite the lack of human\nfeedback for many translations.\n