Category
Instruction & Long Context
Source
Artificial Analysis
evaluation of record
Models covered
449
in our data
Data status
Live
Top score
83.3%
best on record
Top model
Grok 4.3
SpaceXAI
Updated
2026-07-30
last ingest
A benchmark evaluating precise instruction-following generalization on 58 diverse, verifiable out-of-domain constraints that test models' ability to follow specific output requirements.
Leaderboard
Top 20 of 449 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Checks whether a model follows the letter of an instruction across 58 constraints it was not trained on, each one machine-checkable — answer in exactly three sentences, never use a particular word, end on a specific phrase. Higher is better; the score is the share of constraints satisfied, and the quality of the answer is not graded at all, only its compliance. That narrowness is both the limitation and the value: it isolates instruction-following from intelligence, and strong reasoning models routinely lose points here for ignoring a formatting rule they judged unimportant. Because the constraints are held out, it measures generalization rather than the memorized compliance that older instruction benchmarks reward.