AIME 2024

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Models scored

53

evaluated

Modality

text

Category

math

+1 more

Published

2025

arxiv.org

Citations

53

Semantic Scholar

Influential

0

citations

References

44

cited works

Venue

arXiv.org

published in

Abstract

Haoxiang Sun, Yingqian Min, Zhipeng Chen, W. Zhao, et al. (+4)

The rapid advancement of large reasoning models has saturated existing math benchmarks, underscoring the urgent need for more challenging evaluation frameworks. To address this, we introduce OlymMATH, a rigorously curated, Olympiad-level math benchmark comprising 350 problems, each with parallel English and Chinese versions. OlymMATH is the first benchmark to unify dual evaluation paradigms within a single suite: (1) natural language evaluation through OlymMATH-EASY and OlymMATH-HARD, comprising 200 computational problems with numerical answers for objective rule-based assessment, and (2) formal verification through OlymMATH-LEAN, offering 150 problems formalized in Lean 4 for rigorous process-level evaluation. All problems are manually sourced from printed publications to minimize data contamination, verified by experts, and span four core domains. Extensive experiments reveal the benchmark's significant challenge, and our analysis also uncovers consistent performance gaps between languages and identifies cases where models employ heuristic"guessing"rather than rigorous reasoning. To further support community research, we release 582k+ reasoning trajectories, a visualization tool, and expert solutions at https://github.com/RUCAIBox/OlymMATH.

Search

#ModelLabScore
01Grok-3 MinixAI96
02o4-miniOpenAI93
03LongCat-Flash-ThinkingMeituan93
04Grok-3xAI93
05Gemini 2.5 ProGoogle92
06o3OpenAI92
07DeepSeek-R1-0528DeepSeek91
08GLM-4.5Zhipu AI91
09Ministral 3 (14B Reasoning 2512)Mistral AI90
10GLM-4.5-AirZhipu AI89
11Gemini 2.5 FlashGoogle88
12o3-miniOpenAI87
13DeepSeek R1 ZeroDeepSeek87
14DeepSeek R1 Distill Llama 70BDeepSeek87
15MiniMax M1 80KMiniMax86
16o1-proOpenAI86
17Ministral 3 (8B Reasoning 2512)Mistral AI86
18Qwen3 235B A22BAlibaba Cloud / Qwen Team86
19MiniCPM-SALAOpenBMB84
20DeepSeek R1 Distill Qwen 7BDeepSeek83
21DeepSeek R1 Distill Qwen 32BDeepSeek83
22MiniMax M1 40KMiniMax83
23Qwen3 32BAlibaba Cloud / Qwen Team81
24Phi 4 Reasoning PlusMicrosoft81
25Granite 3.3 8B InstructIBM81
26Granite 3.3 8B BaseIBM81
27Qwen3 30B A3BAlibaba Cloud / Qwen Team80
28DeepSeek R1 Distill Llama 8BDeepSeek80
29DeepSeek R1 Distill Qwen 14BDeepSeek80
30Claude 3.7 SonnetAnthropic80
31QwQ-32BAlibaba Cloud / Qwen Team80
32Kimi-k1.5Moonshot AI78
33Min istral 3 (3B Reasoning 2512)Mistral AI78
34Phi 4 ReasoningMicrosoft75
35o1OpenAI74
36Magistral MediumMistral AI74
37Gemini 2.0 Flash ThinkingGoogle73
38LongCat-Flash-LiteMeituan72
39Kimi K2 0905Moonshot AI72
40Magistral Small 2506Mistral AI71
41Kimi K2 InstructMoonshot AI70
42Kimi K2-Instruct-0905Moonshot AI70
43DeepSeek-V3.1DeepSeek66
44DeepSeek-V3 0324DeepSeek59
45DeepSeek R1 Distill Qwen 1.5BDeepSeek53
46QwQ-32B-PreviewAlibaba Cloud / Qwen Team50
47GPT-4.1 miniOpenAI50
48GPT-4.1OpenAI48
49o1-previewOpenAI42
50DeepSeek-V3DeepSeek39
51GPT-4.5OpenAI37
52GPT-4.1 nanoOpenAI29
53GPT-4oOpenAI13

53 of 53 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC