Is human scoring the best criteria for summary evaluation?

Normally, summary quality measures are compared with quality scores produced by human annotators. A higher correlation with human scores is considered to be a fair indicator of a better measure. We discuss observations that cast doubt on this view. We attempt to show a possibility of an alternative indicator. Given a family of measures, we explore a criterion of selecting the best measure not r…

Paper

Similar papers

© 2026 NYSGPT2525 LLC