MATH-500 Benchmark Leaderboard

MathLive

Category

Math

Source

Artificial Analysis

evaluation of record

Models covered

214

in our data

Data status

Live

Top score

99.4%

best on record

Top model

GPT-5

OpenAI

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A 500-problem subset from the MATH dataset, featuring competition-level mathematics across six domains including algebra, geometry, and number theory.

Leaderboard

Top 20 of 214 models we hold a score for.

1GPT-5highOpenAI
99.4%2Grok 3 mini ReasoninghighSpaceXAI
99.2%3o3OpenAI
99.2%4GPT-5mediumOpenAI
99.1%5Claude 4 SonnetReasoningAnthropic
99.1%6Grok 4SpaceXAI
99.0%7o4-minihighOpenAI
98.9%8GPT-5lowOpenAI
98.7%9Gemini 2.5 Pro PreviewMay' 25Google
98.6%10o3-minihighOpenAI
98.5%11Qwen3 235B A22B 2507ReasoningAlibaba
98.4%12Llama Nemotron Super 49B v1.5ReasoningNVIDIA
98.3%13DeepSeek R1 0528May '25DeepSeek
98.3%14Claude 4 OpusReasoningAnthropic
98.2%15Gemini 2.5 FlashReasoningGoogle
98.1%16Gemini 2.5 Flash PreviewReasoningGoogle
98.1%17Gemini 2.5 Pro PreviewMar' 25Google
98.0%18MiniMax M1 80kMiniMax
98.0%19Qwen3 235B A22B 2507 InstructAlibaba
98.0%20GLM-4.5ReasoningZ AI
97.9%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Five hundred competition-level mathematics problems — algebra, geometry, number theory, and three further domains — taken as a representative slice of the much larger MATH dataset. Higher is better, but frontier models now sit above 95 percent, which makes this a saturated benchmark that mostly tells you whether something is broken. Its remaining value is as a regression check rather than a ranking: a model that drops here has a real problem, while two models a point apart are indistinguishable. The underlying dataset is old and heavily reproduced online, so contamination should be assumed rather than ruled out.

© 2026 NYSGPT2525 LLC