AIME 2025

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Models scored

114

evaluated

Modality

text

Category

math

+1 more

Published

2025

arxiv.org

Citations

53

Semantic Scholar

Influential

0

citations

References

44

cited works

Venue

arXiv.org

published in

Abstract

Haoxiang Sun, Yingqian Min, Zhipeng Chen, W. Zhao, et al. (+4)

The rapid advancement of large reasoning models has saturated existing math benchmarks, underscoring the urgent need for more challenging evaluation frameworks. To address this, we introduce OlymMATH, a rigorously curated, Olympiad-level math benchmark comprising 350 problems, each with parallel English and Chinese versions. OlymMATH is the first benchmark to unify dual evaluation paradigms within a single suite: (1) natural language evaluation through OlymMATH-EASY and OlymMATH-HARD, comprising 200 computational problems with numerical answers for objective rule-based assessment, and (2) formal verification through OlymMATH-LEAN, offering 150 problems formalized in Lean 4 for rigorous process-level evaluation. All problems are manually sourced from printed publications to minimize data contamination, verified by experts, and span four core domains. Extensive experiments reveal the benchmark's significant challenge, and our analysis also uncovers consistent performance gaps between languages and identifies cases where models employ heuristic"guessing"rather than rigorous reasoning. To further support community research, we release 582k+ reasoning trajectories, a visualization tool, and expert solutions at https://github.com/RUCAIBox/OlymMATH.

Search

#ModelLabScore
01Grok-4 HeavyxAI100
02Kimi K2-Thinking-0905Moonshot AI100
03Gemini 3 ProGoogle100
04GPT-5.2OpenAI100
05GPT-5.2 ProOpenAI100
06Claude Opus 4.6Anthropic100
07Gemini 3 FlashGoogle100
08GPT-5.1 HighOpenAI100
09LongCat-Flash-Thinking-2601Meituan100
10Nemotron 3 Nano (30B A3B)NVIDIA99
11GPT OSS 20B HighOpenAI99
12GPT-5.1 MediumOpenAI98
13Seed 2.0 ProByteDance98
14Step-3.5-FlashStepFun97
15MAI-Thinking-1Microsoft97
16Sarvam-30BSarvam AI97
17Sarvam-105BSarvam AI97
18GPT-5.1 Codex HighOpenAI97
19Kimi K2.5Moonshot AI96
20DeepSeek-V3.2-SpecialeDeepSeek96
21GLM-4.7Zhipu AI96
22GPT-5OpenAI95
23GPT-5 HighOpenAI95
24MiMo-V2-FlashXiaomi94
25GPT-5.1 ThinkingOpenAI94
26GPT-5.1 InstantOpenAI94
27GPT-5.1OpenAI94
28GLM-4.6Zhipu AI94
29Grok-3xAI93
30DeepSeek-V3.2DeepSeek93
31DeepSeek-V3.2 (Thinking)DeepSeek93
32Seed 2.0 LiteByteDance93
33K-EXAONE-236B-A23BLG AI Research93
34o4-miniOpenAI93
35GPT OSS 120B HighOpenAI93
36Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team92
37Nova 2 ProAmazon92
38Nova 2 OmniAmazon92
39Grok 4 FastxAI92
40Grok-4xAI92
41GLM-4.7-FlashZhipu AI92
42Mercury 2Inception91
43GPT-5 miniOpenAI91
44Nova 2 LiteAmazon91
45Grok-3 MinixAI91
46LongCat-Flash-ThinkingMeituan91
47Nemotron 3 Super (120B A12B)NVIDIA90
48Command A+Cohere90
49Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team90
50DeepSeek-V3.2-ExpDeepSeek89
51GPT-5 MediumOpenAI89
52Gemini 2.5 Pro Preview 06-05Google88
53Qwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team88
54Step3-VL-10BStepFun88
55DeepSeek-R1-0528DeepSeek88
56Claude Sonnet 4.5Anthropic87
57ERNIE 5.0Baidu87
58o3OpenAI86
59Mistral Medium 3.5Mistral AI86
60GPT-5 nanoOpenAI85

60 of 114 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC