A critical analysis of metrics used for measuring progress in artificial intelligence

Comparing model performances on benchmark datasets is an integral part of\nmeasuring and driving progress in artificial intelligence. A model's\nperformance on a benchmark dataset is commonly assessed based on a single or a\nsmall set of performance metrics. While this enables quick comparisons, it may\nentail the risk of inadequately reflecting model performance if the metric does\nnot sufficiently cover all performance characteristics. It is unknown to what\nextent this might impact benchmarking efforts.\n To address this question, we analysed the current landscape of performance\nmetrics based on data covering 3867 machine learning model performance results\nfrom the open repository 'Papers with Code'. Our results suggest that the large\nmajority of metrics currently used have properties that may result in an\ninadequate reflection of a models' performance. While alternative metrics that\naddress problematic properties have been proposed, they are currently rarely\nused.\n Furthermore, we describe ambiguities in reported metrics, which may lead to\ndifficulties in interpreting and comparing model performances.\n

Paper

References (65)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC