MATH

Measuring Mathematical Problem Solving With the MATH Dataset

Models scored

71

evaluated

Modality

text

Category

math

+1 more

Published

2021

arxiv.org

Citations

5,501

Semantic Scholar

Influential

1,246

citations

References

71

cited works

Venue

NeurIPS Datasets and Benchmarks

published in

Abstract

Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, et al. (+4)

Many intellectual endeavors require mathematical problem solving, but this skill remains beyond the capabilities of computers. To measure this ability in machine learning models, we introduce MATH, a new dataset of 12,500 challenging competition mathematics problems. Each problem in MATH has a full step-by-step solution which can be used to teach models to generate answer derivations and explanations. To facilitate future research and increase accuracy on MATH, we also contribute a large auxiliary pretraining dataset which helps teach models the fundamentals of mathematics. Even though we are able to increase accuracy on MATH, our results show that accuracy remains relatively low, even with enormous Transformer models. Moreover, we find that simply increasing budgets and model parameter counts will be impractical for achieving strong mathematical reasoning if scaling trends continue. While scaling Transformers is automatically solving most other text-based tasks, scaling is not currently solving MATH. To have more traction on mathematical problem solving we will likely need new algorithmic advancements from the broader research community.

Search

#ModelLabScore
01o3-miniOpenAI98
02o1OpenAI96
03Mistral Large 3Mistral AI90
04MiniStral 3 (14B Instruct 2512)Mistral AI90
05Gemini 2.0 FlashGoogle90
06Kimi K2 0905Moonshot AI89
07Gemma 3 27BGoogle89
08Ministral 3 (8B Instruct 2512)Mistral AI88
09Gemini 2.0 Flash-LiteGoogle87
10Gemini 1.5 ProGoogle87
11MiMo-V2.5-ProXiaomi86
12o1-previewOpenAI86
13GPT-5OpenAI85
14Gemma 3 12BGoogle84
15Qwen2.5 72B InstructAlibaba Cloud / Qwen Team83
16Qwen2.5 32B InstructAlibaba Cloud / Qwen Team83
17Ministral 3 (3B Instruct 2512)Mistral AI83
18Qwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team82
19Phi 4Microsoft80
20Qwen2.5 14B InstructAlibaba Cloud / Qwen Team80
21Claude 3.5 SonnetAnthropic78
22Gemini 1.5 FlashGoogle78
23Llama 3.3 70B InstructMeta77
24Nova ProAmazon77
25GPT-4oOpenAI77
26Grok-2xAI76
27Gemma 3 4BGoogle76
28Qwen2.5 7B InstructAlibaba Cloud / Qwen Team76
29DeepSeek-V2.5DeepSeek75
30Llama 3.1 405B InstructMeta74
31Nova LiteAmazon73
32Grok-2 minixAI73
33GPT-4 TurboOpenAI73
34Qwen3 235B A22BAlibaba Cloud / Qwen Team72
35Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team72
36Claude 3.5 SonnetAnthropic71
37Mistral Small 3 24B InstructMistral AI71
38GPT-4o miniOpenAI70
39Kimi K2 BaseMoonshot AI70
40Mistral Small 3.2 24B InstructMistral AI69
41Claude 3.5 HaikuAnthropic69
42Nova MicroAmazon69
43Mistral Small 3.1 24B InstructMistral AI69
44Llama 3.2 90B InstructMeta68
45Phi 4 MiniMicrosoft64
46Llama 4 MaverickMeta61
47Claude 3 OpusAnthropic60
48Qwen2 72B InstructAlibaba Cloud / Qwen Team60
49Phi-3.5-MoE-instructMicrosoft60
50Gemini 1.5 Flash 8BGoogle59
51Qwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team57
52Ministral 8B InstructMistral AI55
53Llama 3.2 11B InstructMeta52
54Grok-1.5xAI51
55Llama 4 ScoutMeta50
56Qwen2 7B InstructAlibaba Cloud / Qwen Team50
57Phi-3.5-mini-instructMicrosoft49
58Pixtral-12BMistral AI48
59Gemma 3 1BGoogle48
60Llama 3.2 3B InstructMeta48

60 of 71 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC