Benchmarks for tool-using LLM agents are undergoing a visible design shift: away from pooled single scores and toward explicit regimes, explicit admissibility, and checkable closure. This paper argues that the shift is not merely incremental evaluation engineering; it is operational convergence toward the reliability unit implied by Large Language Fields (LLFs): a field-level contract layer above the model where meaning, evidence, and “done” are governed by reusable, versioned rules. We formalize the convergence as a move from impression-based scoring to contract-based verification. Subfields become declared regime slices with distinct closure predicates, and progress becomes portable only when supported by receipts that an independent verifier can validate under a declared verification budget. We further show how this contract lens aligns with collapse-style diagnosis: pooled means compress heterogeneous regimes into a single number that can appear stable while semantic integrity fails in the tails. We connect emerging benchmark patterns — interactive research tasks, controllable extreme context growth, state-diff outcome contracts, and construct-before-prove formal tasks — to a unified LLF reading: benchmarks are increasingly instrumenting the field contract rather than merely sampling model outputs. We then state falsifiable predictions: regime-explicit benchmarks will predict operational failures better than pooled means; worst-subfield performance, dispersion, and regime-flip maps will dominate mean scores as predictors of governance outcomes; and receipt completeness under audit sampling within budget will explain a large fraction of false closure in agent deployments.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex