HMMT 2025

Towards Self-Verifiable Mathematical Reasoning

Models scored

33

evaluated

Modality

text

Category

math

Published

2025

web.mit.edu

Citations

45

Semantic Scholar

Influential

6

citations

References

20

cited works

Venue

arXiv.org

published in

Abstract

Zhihong Shao, Yu-Wei Luo, Chengda Lu, Z. Ren, et al. (+5)

Large language models have made significant progress in mathematical reasoning, which serves as an important testbed for AI and could impact scientific research if further advanced. By scaling reasoning with reinforcement learning that rewards correct final answers, LLMs have improved from poor performance to saturating quantitative reasoning competitions like AIME and HMMT in one year. However, this approach faces fundamental limitations. Pursuing higher final answer accuracy doesn't address a key issue: correct answers don't guarantee correct reasoning. Moreover, many mathematical tasks like theorem proving require rigorous step-by-step derivation rather than numerical answers, making final answer rewards inapplicable. To push the limits of deep reasoning, we believe it is necessary to verify the comprehensiveness and rigor of mathematical reasoning. Self-verification is particularly important for scaling test-time compute, especially for open problems without known solutions. Towards self-verifiable mathematical reasoning, we investigate how to train an accurate and faithful LLM-based verifier for theorem proving. We then train a proof generator using the verifier as the reward model, and incentivize the generator to identify and resolve as many issues as possible in their own proofs before finalizing them. To maintain the generation-verification gap as the generator becomes stronger, we propose to scale verification compute to automatically label new hard-to-verify proofs, creating training data to further improve the verifier. Our resulting model, DeepSeekMath-V2, demonstrates strong theorem-proving capabilities, achieving gold-level scores on IMO 2025 and CMO 2024 and a near-perfect 118/120 on Putnam 2024 with scaled test-time compute.

Search

#ModelLabScore
01GPT-5.2 ProOpenAI100
02GPT-5.2OpenAI99
03DeepSeek-V3.2-SpecialeDeepSeek99
04Kimi K2-Thinking-0905Moonshot AI98
05Qwen3.6 PlusAlibaba Cloud / Qwen Team97
06Kimi K2.5Moonshot AI95
07Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team95
08Nemotron 3 Super (120B A12B)NVIDIA95
09GLM-5.2Zhipu AI94
10GLM-5.1Zhipu AI94
11Qwen3.6-27BAlibaba Cloud / Qwen Team94
12GPT-5OpenAI93
13Grok 4 FastxAI93
14Qwen3.5-27BAlibaba Cloud / Qwen Team92
15Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team91
16Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team91
17DeepSeek-V3.2DeepSeek90
18DeepSeek-V3.2 (Thinking)DeepSeek90
19Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team89
20GPT-5 miniOpenAI88
21Sarvam-105BSarvam AI86
22MiMo-V2-FlashXiaomi84
23DeepSeek-V3.2-ExpDeepSeek84
24Qwen3.5-9BAlibaba Cloud / Qwen Team83
25DeepSeek-R1-0528DeepSeek79
26GPT-5 nanoOpenAI76
27Qwen3.5-4BAlibaba Cloud / Qwen Team74
28Sarvam-30BSarvam AI73
29Kimi K2-Instruct-0905Moonshot AI39
30Kimi K2 InstructMoonshot AI39
31GPT-4.1 miniOpenAI35
32DeepSeek-V3.1DeepSeek34
33GPT-4.1OpenAI29

33 of 33 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC