Summary
This position paper proposes a new discipline: "model metrology", aimed at generating new benchmarks for Language Models (LM) evaluation that overcome the limitations of current benchmarks (saturation, memorisation, non-constrained scope, etc.). First, the paper analyses current limitations of current benchmarks, then proposes a number of desired features for such a new generation of benchmarks: constrained, dynamic, and plug-and-play. Finally, it analyses the characteristics of the proposed "model metrology" discipline and proposes a roadmap to achieve such a vision.
The topic is a relevant and interesting one. Limitations of current benchmarks for LLM evaluation are well known, and there are active research efforts in the direction of mitigating them, to which this work could contribute. The paper contains interesting ideas (the idea of model metrology as a new discipline is appealing) and can be the basis for stimulating a debate in that direction.
However, the paper has some flaws in my view. In general, there are a number of good ideas, but some of them not enough elaborated, some of them are bit vague (e.g., section 4.3) and some others not so novel, such as the "constrained" feature (task-based and extrinsic evaluations have been around for decades). In contrast, the "dynamic" aspect is a more contemporary need in NLP. However, I miss the "human in the loop" aspect as well.
In the paper, there are too subjective claims such as "[OpenAI, Antrophic, Google] still cling to the idea that general intelligence is quantifiable by these benchmarks", which are not sufficiently supported. (Do these vendors really think so? If so, citations would be welcome).
As an example of the limitations of current benchmarks, the Air Canada Chatbot incident is presented. In fact, as stated by the authors, is "implausible that this failure case could have been predicted through generalized benchmark results". However, this does not contradict the value of such benchmarks to allow testing (and improving) the model along several specific dimensions, before integrating them in a final system. The tests of the system as a whole (model + policies specification + UI, etc.) in the Air Canada case were obviously incomplete, but this is possibly because good test engineering practices were not followed, not invalidating the current LLM benchmarks per se.
From a stylistic point of view, the writing style is a bit wordy at times and difficult to follow, lacking the conciseness of academic writing. Some grammar constructions are a bit weird, for instance in section 3: "One reason [for what?] is that static benchmarks of that form [which form?] are easily memorized, […] “, or in "outlook and roadblocks to this discussed in §4.2", where the auxiliary verb is missing. There are also many contractions along the whole text (aren't, it's, can't), which is unusual in academic written English.
Despite the extensive references section, there are some omissions that could have complemented the author's analysis, such as: Tedeschi et al. "What’s the Meaning of Superhuman Performance in Today’s NLU?", ACL 2023 (a critical analysis of superhuman claims derived from using traditional benchmarks, with ideas to incentivize benchmark creators to design more solid and transparent benchmarks).