MMLU-Pro
A More Robust and Challenging Multi-Task Language Understanding Benchmark
Models scored
129
evaluated
Modality
text
Category
finance
+6 more
Published
2024
arxiv.org
Citations
1,729
Semantic Scholar
Influential
197
citations
References
56
cited works
Venue
Neural Information Processing Systems
published in
Abstract
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, et al. (+13)
In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning across diverse domains. However, as models continue to improve, their performance on these benchmarks has begun to plateau, making it increasingly difficult to discern differences in model capabilities. This paper introduces MMLU-Pro, an enhanced dataset designed to extend the mostly knowledge-driven MMLU benchmark by integrating more challenging, reasoning-focused questions and expanding the choice set from four to ten options. Additionally, MMLU-Pro eliminates the trivial and noisy questions in MMLU. Our experimental results show that MMLU-Pro not only raises the challenge, causing a significant drop in accuracy by 16% to 33% compared to MMLU but also demonstrates greater stability under varying prompts. With 24 different prompt styles tested, the sensitivity of model scores to prompt variations decreased from 4-5% in MMLU to just 2% in MMLU-Pro. Additionally, we found that models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering, which is in stark contrast to the findings on the original MMLU, indicating that MMLU-Pro includes more complex reasoning questions. Our assessments confirm that MMLU-Pro is a more discriminative benchmark to better track progress in the field.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Qwen3.7 Max | Alibaba Cloud / Qwen Team | 90 |
| 02 | Qwen3.7-Plus | Alibaba Cloud / Qwen Team | 89 |
| 03 | Qwen3.6 Plus | Alibaba Cloud / Qwen Team | 89 |
| 04 | MiniMax M2.1 | MiniMax | 88 |
| 05 | Qwen3.5-397B-A17B | Alibaba Cloud / Qwen Team | 88 |
| 06 | DeepSeek-V4-Pro-Max | DeepSeek | 88 |
| 07 | Kimi K2.5 | Moonshot AI | 87 |
| 08 | ERNIE 5.0 | Baidu | 87 |
| 09 | Nemotron 3 Ultra (550B A55B) | NVIDIA | 87 |
| 10 | Qwen3.5-122B-A10B | Alibaba Cloud / Qwen Team | 87 |
| 11 | DeepSeek-V4-Flash-Max | DeepSeek | 86 |
| 12 | Qwen3.6-27B | Alibaba Cloud / Qwen Team | 86 |
| 13 | Qwen3.5-27B | Alibaba Cloud / Qwen Team | 86 |
| 14 | Qwen3.5-35B-A3B | Alibaba Cloud / Qwen Team | 85 |
| 15 | Gemma 4 31B | 85 | |
| 16 | Qwen3.6-35B-A3B | Alibaba Cloud / Qwen Team | 85 |
| 17 | DeepSeek-R1-0528 | DeepSeek | 85 |
| 18 | DeepSeek-V3.2-Exp | DeepSeek | 85 |
| 19 | DeepSeek-V3.2 | DeepSeek | 85 |
| 20 | DeepSeek-V3.2 (Thinking) | DeepSeek | 85 |
| 21 | MAI-Thinking-1 | Microsoft | 85 |
| 22 | MiMo-V2-Flash | Xiaomi | 85 |
| 23 | GLM-4.5 | Zhipu AI | 85 |
| 24 | Kimi K2-Thinking-0905 | Moonshot AI | 85 |
| 25 | Qwen3-235B-A22B-Thinking-2507 | Alibaba Cloud / Qwen Team | 84 |
| 26 | GLM-4.7 | Zhipu AI | 84 |
| 27 | K-EXAONE-236B-A23B | LG AI Research | 84 |
| 28 | Qwen3 VL 235B A22B Thinking | Alibaba Cloud / Qwen Team | 84 |
| 29 | Nemotron 3 Super (120B A12B) | NVIDIA | 84 |
| 30 | DeepSeek-V3.1 | DeepSeek | 84 |
| 31 | Qwen3-235B-A22B-Instruct-2507 | Alibaba Cloud / Qwen Team | 83 |
| 32 | Qwen3-Next-80B-A3B-Thinking | Alibaba Cloud / Qwen Team | 83 |
| 33 | LongCat-Flash-Chat | Meituan | 83 |
| 34 | Gemma 4 26B-A4B | 83 | |
| 35 | LongCat-Flash-Thinking | Meituan | 83 |
| 36 | Kimi K2 0905 | Moonshot AI | 83 |
| 37 | Qwen3.5-9B | Alibaba Cloud / Qwen Team | 83 |
| 38 | Qwen3 VL 32B Thinking | Alibaba Cloud / Qwen Team | 82 |
| 39 | MiniMax M2 | MiniMax | 82 |
| 40 | Qwen3 VL 235B A22B Instruct | Alibaba Cloud / Qwen Team | 82 |
| 41 | Sarvam-105B | Sarvam AI | 82 |
| 42 | Nova 2 Pro | Amazon | 82 |
| 43 | GLM-4.5-Air | Zhipu AI | 81 |
| 44 | DeepSeek-V3 0324 | DeepSeek | 81 |
| 45 | Kimi K2 Instruct | Moonshot AI | 81 |
| 46 | Kimi K2-Instruct-0905 | Moonshot AI | 81 |
| 47 | MiniMax M1 80K | MiniMax | 81 |
| 48 | Nova 2 Lite | Amazon | 81 |
| 49 | Nova 2 Omni | Amazon | 81 |
| 50 | GPT OSS 120B High | OpenAI | 81 |
| 51 | Qwen3-Next-80B-A3B-Instruct | Alibaba Cloud / Qwen Team | 81 |
| 52 | MiniMax M1 40K | MiniMax | 81 |
| 53 | Qwen3 VL 30B A3B Thinking | Alibaba Cloud / Qwen Team | 81 |
| 54 | Llama 4 Maverick | Meta | 81 |
| 55 | Sarvam-30B | Sarvam AI | 80 |
| 56 | Qwen3.5-4B | Alibaba Cloud / Qwen Team | 79 |
| 57 | Qwen3 VL 32B Instruct | Alibaba Cloud / Qwen Team | 79 |
| 58 | Nemotron 3 Nano (30B A3B) | NVIDIA | 78 |
| 59 | LongCat-Flash-Lite | Meituan | 78 |
| 60 | Mistral Small 4 | Mistral AI | 78 |
60 of 129 models · score normalized 0–100 where available