Category
Math
Source
Artificial Analysis
evaluation of record
Models covered
214
in our data
Data status
Live
Top score
99.4%
best on record
Top model
GPT-5
OpenAI
Updated
2026-07-30
last ingest
A 500-problem subset from the MATH dataset, featuring competition-level mathematics across six domains including algebra, geometry, and number theory.
Leaderboard
Top 20 of 214 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Five hundred competition-level mathematics problems — algebra, geometry, number theory, and three further domains — taken as a representative slice of the much larger MATH dataset. Higher is better, but frontier models now sit above 95 percent, which makes this a saturated benchmark that mostly tells you whether something is broken. Its remaining value is as a regression check rather than a ranking: a model that drops here has a real problem, while two models a point apart are indistinguishable. The underlying dataset is old and heavily reproduced online, so contamination should be assumed rather than ruled out.