HMMT25

Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B

Models scored

25

evaluated

Modality

text

Category

math

Published

2025

web.mit.edu

Citations

6

Semantic Scholar

Influential

0

citations

References

28

cited works

Venue

arXiv.org

published in

Abstract

Sen Xu, Yi Zhou, Wei Wang, Jixin Min, et al. (+6)

Challenging the prevailing consensus that small models inherently lack robust reasoning, this report introduces VibeThinker-1.5B, a 1.5B-parameter dense model developed via our Spectrum-to-Signal Principle (SSP). This challenges the prevailing approach of scaling model parameters to enhance capabilities, as seen in models like DeepSeek R1 (671B) and Kimi k2 (>1T). The SSP framework first employs a Two-Stage Diversity-Exploring Distillation (SFT) to generate a broad spectrum of solutions, followed by MaxEnt-Guided Policy Optimization (RL) to amplify the correct signal. With a total training cost of only $7,800, VibeThinker-1.5B demonstrates superior reasoning capabilities compared to closed-source models like Magistral Medium and Claude Opus 4, and performs on par with open-source models like GPT OSS-20B Medium. Remarkably, it surpasses the 400x larger DeepSeek R1 on three math benchmarks: AIME24 (80.3 vs. 79.8), AIME25 (74.4 vs. 70.0), and HMMT25 (50.4 vs. 41.7). This is a substantial improvement over its base model (6.7, 4.3, and 0.6, respectively). On LiveCodeBench V6, it scores 51.1, outperforming Magistral Medium's 50.3 and its base model's 0.0. These findings demonstrate that small models can achieve reasoning capabilities comparable to large models, drastically reducing training and inference costs and thereby democratizing advanced AI research.

Search

#ModelLabScore
01Grok-4 HeavyxAI97
02Qwen3.6 PlusAlibaba Cloud / Qwen Team95
03Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team93
04Qwen3.6-27BAlibaba Cloud / Qwen Team91
05Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team90
06Grok-4xAI90
07Qwen3.5-27BAlibaba Cloud / Qwen Team90
08Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team89
09Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team89
10Sarvam-105BSarvam AI86
11Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team84
12Qwen3.5-9BAlibaba Cloud / Qwen Team83
13Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team77
14Qwen3.5-4BAlibaba Cloud / Qwen Team77
15Sarvam-30BSarvam AI74
16Qwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team74
17Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team68
18Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team61
19Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team57
20Qwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team55
21Qwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team54
22Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team53
23Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team51
24Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team33
25Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team31

25 of 25 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC