Automatic evaluation of language generation systems is a well-studied problem\nin Natural Language Processing. While novel metrics are proposed every year, a\nfew popular metrics remain as the de facto metrics to evaluate tasks such as\nimage captioning and machine translation, despite their known limitations. This\nis partly due to ease of use, and partly because researchers expect to see them\nand know how to interpret them. In this paper, we urge the community for more\ncareful consideration of how they automatically evaluate their models by\ndemonstrating important failure cases on multiple datasets, language pairs and\ntasks. Our experiments show that metrics (i) usually prefer system outputs to\nhuman-authored texts, (ii) can be insensitive to correct translations of rare\nwords, (iii) can yield surprisingly high scores when given a single sentence as\nsystem output for the entire test set.\n
Paper
References (27)
Scroll for more · 15 remaining