GSM8k

Training Verifiers to Solve Math Word Problems

Models scored

48

evaluated

Modality

text

Category

math

+1 more

Published

2021

arxiv.org

Citations

9,839

Semantic Scholar

Influential

2,175

citations

References

31

cited works

Venue

arXiv.org

published in

Abstract

K. Cobbe, Vineet Kosaraju, Mo Bavarian, Mark Chen, et al. (+8)

State-of-the-art language models can match human performance on many tasks, but they still struggle to robustly perform multi-step mathematical reasoning. To diagnose the failures of current models and support research, we introduce GSM8K, a dataset of 8.5K high quality linguistically diverse grade school math word problems. We find that even the largest transformer models fail to achieve high test performance, despite the conceptual simplicity of this problem distribution. To increase performance, we propose training verifiers to judge the correctness of model completions. At test time, we generate many candidate solutions and select the one ranked highest by the verifier. We demonstrate that verification significantly improves performance on GSM8K, and we provide strong empirical evidence that verification scales more effectively with increased data than a finetuning baseline.

Search

#ModelLabScore
01MiMo-V2.5-ProXiaomi100
02Kimi K2 InstructMoonshot AI97
03o1OpenAI97
04GPT-4.5OpenAI97
05Llama 3.1 405B InstructMeta97
06Claude 3.5 SonnetAnthropic96
07Claude 3.5 SonnetAnthropic96
08Gemma 3 27BGoogle96
09Qwen2.5 32B InstructAlibaba Cloud / Qwen Team96
10Qwen2.5 72B InstructAlibaba Cloud / Qwen Team96
11DeepSeek-V2.5DeepSeek95
12Claude 3 OpusAnthropic95
13Nova ProAmazon95
14Qwen2.5 14B InstructAlibaba Cloud / Qwen Team95
15Nova LiteAmazon95
16Gemma 3 12BGoogle94
17Qwen3 235B A22BAlibaba Cloud / Qwen Team94
18Mistral Large 2Mistral AI93
19Claude 3 SonnetAnthropic92
20Nova MicroAmazon92
21Kimi K2 BaseMoonshot AI92
22Qwen2.5 7B InstructAlibaba Cloud / Qwen Team92
23Llama 3.1 Nemotron 70B InstructNVIDIA91
24Qwen2 72B InstructAlibaba Cloud / Qwen Team91
25Qwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team91
26Gemini 1.5 ProGoogle91
27Grok-1.5xAI90
28Gemma 3 4BGoogle89
29Claude 3 HaikuAnthropic89
30Phi-3.5-MoE-instructMicrosoft89
31Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team89
32Phi 4 MiniMicrosoft89
33Jamba 1.5 LargeAI21 Labs87
34Phi-3.5-mini-instructMicrosoft86
35Gemini 1.5 FlashGoogle86
36Qwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team84
37Qwen2 7B InstructAlibaba Cloud / Qwen Team82
38Granite 3.3 8B InstructIBM81
39Mistral Small 3 24B BaseMistral AI81
40Llama 3.2 3B InstructMeta78
41Jamba 1.5 MiniAI21 Labs76
42Gemma 2 27BGoogle74
43Command R+Cohere71
44IBM Granite 4.0 Tiny PreviewIBM70
45Gemma 2 9BGoogle69
46Gemma 3 1BGoogle63
47Granite 3.3 8B BaseIBM59
48ERNIE 4.5Baidu25

48 of 48 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC