Category
Knowledge & Reasoning
Source
Artificial Analysis
evaluation of record
Models covered
547
in our data
Data status
Live
Top score
53.3%
best on record
Top model
Claude Fable 5
Anthropic
Updated
2026-07-30
last ingest
A frontier-level benchmark with 2,500 expert-vetted questions across mathematics, sciences, and humanities, designed to be the final closed-ended academic evaluation.
Leaderboard
Top 20 of 547 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Two and a half thousand questions written and vetted by specialists to sit at the edge of what an expert in the field could answer, spanning mathematics, the sciences, and the humanities. Higher is better; the frontier is now around half, up from single digits two years ago, so 10-point gaps here still separate genuine capability tiers rather than noise. Scores depend heavily on whether a model was allowed tools and search — a tools-enabled run and a closed-book run are different measurements that get quoted interchangeably, so always check which one a number is. It was built to be the last closed-ended academic exam worth running, which means it will saturate, probably sooner than its authors intended.