AIME 2025 Benchmark Leaderboard

MathLive

Category

Math

Source

Artificial Analysis

evaluation of record

Models covered

269

in our data

Data status

Live

Top score

99.0%

best on record

Top model

GPT-5.2

OpenAI

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

All 30 problems from the 2025 American Invitational Mathematics Examination, testing olympiad-level mathematical reasoning with integer answers from 000-999.

Leaderboard

Top 20 of 269 models we hold a score for.

1GPT-5.2xhighOpenAI
99.0%2GPT-5 CodexhighOpenAI
98.7%3Gemini 3 Flash PreviewReasoningGoogle
97.0%4DeepSeek V3.2 SpecialeDeepSeek
96.7%5GPT-5.2mediumOpenAI
96.7%6MiMo-V2-FlashReasoningXiaomi
96.3%7GPT-5.1 CodexhighOpenAI
95.7%8Gemini 3 Pro PreviewhighGoogle
95.7%9GLM-4.7ReasoningZ AI
95.0%10KAT-Coder-Pro V1KwaiKAT
94.7%11Kimi K2 ThinkingKimi
94.7%12GPT-5highOpenAI
94.3%13Nova 2.0 LitehighAmazon
94.3%14GPT-5.1highOpenAI
94.0%15gpt-oss-120bhighOpenAI
93.4%16Grok 4SpaceXAI
92.7%17DeepSeek V3.2ReasoningDeepSeek
92.0%18GPT-5mediumOpenAI
91.7%19GPT-5.1 Codex minihighOpenAI
91.7%20Claude Opus 4.5ReasoningAnthropic
91.3%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

All 30 problems from the 2025 American Invitational Mathematics Examination, the qualifying round for the US Mathematical Olympiad, where every answer is an integer from 000 to 999 and can therefore be marked exactly. Higher is better — but with only 30 problems each one is worth 3.3 points, so a two-problem difference looks like a 7-point gap. Labs usually report an average over many attempts for exactly that reason, and a single-run score is not comparable with an averaged one. Using the 2025 paper keeps it fresh relative to training cut-offs, but that advantage expires as the questions and their solutions spread online.

© 2026 NYSGPT2525 LLC