MATH-500

Measuring Mathematical Problem Solving With the MATH Dataset

Models scored

32

evaluated

Modality

text

Category

math

+1 more

Published

2021

arxiv.org

Citations

5,501

Semantic Scholar

Influential

1,246

citations

References

71

cited works

Venue

NeurIPS Datasets and Benchmarks

published in

Abstract

Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, et al. (+4)

Many intellectual endeavors require mathematical problem solving, but this skill remains beyond the capabilities of computers. To measure this ability in machine learning models, we introduce MATH, a new dataset of 12,500 challenging competition mathematics problems. Each problem in MATH has a full step-by-step solution which can be used to teach models to generate answer derivations and explanations. To facilitate future research and increase accuracy on MATH, we also contribute a large auxiliary pretraining dataset which helps teach models the fundamentals of mathematics. Even though we are able to increase accuracy on MATH, our results show that accuracy remains relatively low, even with enormous Transformer models. Moreover, we find that simply increasing budgets and model parameter counts will be impractical for achieving strong mathematical reasoning if scaling trends continue. While scaling Transformers is automatically solving most other text-based tasks, scaling is not currently solving MATH. To have more traction on mathematical problem solving we will likely need new algorithmic advancements from the broader research community.

Search

#ModelLabScore
01LongCat-Flash-ThinkingMeituan99
02Sarvam-105BSarvam AI99
03GLM-4.5Zhipu AI98
04GLM-4.5-AirZhipu AI98
05Nemotron Nano 9B v2NVIDIA98
06Kimi K2-Instruct-0905Moonshot AI97
07Kimi K2 InstructMoonshot AI97
08Llama 3.1 Nemotron Ultra 253B v1NVIDIA97
09Sarvam-30BSarvam AI97
10MiniMax M1 80KMiniMax97
11LongCat-Flash-LiteMeituan97
12Llama-3.3 Nemotron Super 49B v1NVIDIA97
13LongCat-Flash-ChatMeituan96
14Kimi-k1.5Moonshot AI96
15Claude 3.7 SonnetAnthropic96
16MiniMax M1 40KMiniMax96
17DeepSeek R1 ZeroDeepSeek96
18Llama 3.1 Nemotron Nano 8B V1NVIDIA95
19Phi 4 Mini ReasoningMicrosoft95
20DeepSeek R1 Distill Llama 70BDeepSeek95
21DeepSeek R1 Distill Qwen 32BDeepSeek94
22DeepSeek-V3 0324DeepSeek94
23DeepSeek R1 Distill Qwen 14BDeepSeek94
24DeepSeek R1 Distill Qwen 7BDeepSeek93
25QwQ-32BAlibaba Cloud / Qwen Team91
26QwQ-32B-PreviewAlibaba Cloud / Qwen Team91
27DeepSeek-V3DeepSeek90
28o1-miniOpenAI90
29DeepSeek R1 Distill Llama 8BDeepSeek89
30DeepSeek R1 Distill Qwen 1.5BDeepSeek84
31Granite 3.3 8B BaseIBM69
32Granite 3.3 8B InstructIBM69

32 of 32 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC