ACUTE-EVAL: Improved Dialogue Evaluation with Optimized Questions and Multi-turn Comparisons

While dialogue remains an important end-goal of natural language research,\nthe difficulty of evaluation is an oft-quoted reason why it remains troublesome\nto make real progress towards its solution. Evaluation difficulties are\nactually two-fold: not only do automatic metrics not correlate well with human\njudgments, but also human judgments themselves are in fact difficult to\nmeasure. The two most used human judgment tests, single-turn pairwise\nevaluation and multi-turn Likert scores, both have serious flaws as we discuss\nin this work.\n We instead provide a novel procedure involving comparing two full dialogues,\nwhere a human judge is asked to pay attention to only one speaker within each,\nand make a pairwise judgment. The questions themselves are optimized to\nmaximize the robustness of judgments across different annotators, resulting in\nbetter tests. We also show how these tests work in self-play model chat setups,\nresulting in faster, cheaper tests. We hope these tests become the de facto\nstandard, and will release open-source code to that end.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC