Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale

Automatic evaluation of language generation systems is a well-studied problem\nin Natural Language Processing. While novel metrics are proposed every year, a\nfew popular metrics remain as the de facto metrics to evaluate tasks such as\nimage captioning and machine translation, despite their known limitations. This\nis partly due to ease of use, and partly because researchers expect to see them\nand know how to interpret them. In this paper, we urge the community for more\ncareful consideration of how they automatically evaluate their models by\ndemonstrating important failure cases on multiple datasets, language pairs and\ntasks. Our experiments show that metrics (i) usually prefer system outputs to\nhuman-authored texts, (ii) can be insensitive to correct translations of rare\nwords, (iii) can yield surprisingly high scores when given a single sentence as\nsystem output for the entire test set.\n

Paper

References (27)

Scroll for more · 15 remaining

Similar papers

© 2026 NYSGPT2525 LLC