MMLU-Pro

A More Robust and Challenging Multi-Task Language Understanding Benchmark

Models scored

129

evaluated

Modality

text

Category

finance

+6 more

Published

2024

arxiv.org

Citations

1,729

Semantic Scholar

Influential

197

citations

References

56

cited works

Venue

Neural Information Processing Systems

published in

Abstract

Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, et al. (+13)

In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning across diverse domains. However, as models continue to improve, their performance on these benchmarks has begun to plateau, making it increasingly difficult to discern differences in model capabilities. This paper introduces MMLU-Pro, an enhanced dataset designed to extend the mostly knowledge-driven MMLU benchmark by integrating more challenging, reasoning-focused questions and expanding the choice set from four to ten options. Additionally, MMLU-Pro eliminates the trivial and noisy questions in MMLU. Our experimental results show that MMLU-Pro not only raises the challenge, causing a significant drop in accuracy by 16% to 33% compared to MMLU but also demonstrates greater stability under varying prompts. With 24 different prompt styles tested, the sensitivity of model scores to prompt variations decreased from 4-5% in MMLU to just 2% in MMLU-Pro. Additionally, we found that models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering, which is in stark contrast to the findings on the original MMLU, indicating that MMLU-Pro includes more complex reasoning questions. Our assessments confirm that MMLU-Pro is a more discriminative benchmark to better track progress in the field.

financegeneralhealthcarelanguagelegalmathreasoning

Search

#ModelLabScore
01Qwen3.7 MaxAlibaba Cloud / Qwen Team90
02Qwen3.7-PlusAlibaba Cloud / Qwen Team89
03Qwen3.6 PlusAlibaba Cloud / Qwen Team89
04MiniMax M2.1MiniMax88
05Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team88
06DeepSeek-V4-Pro-MaxDeepSeek88
07Kimi K2.5Moonshot AI87
08ERNIE 5.0Baidu87
09Nemotron 3 Ultra (550B A55B)NVIDIA87
10Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team87
11DeepSeek-V4-Flash-MaxDeepSeek86
12Qwen3.6-27BAlibaba Cloud / Qwen Team86
13Qwen3.5-27BAlibaba Cloud / Qwen Team86
14Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team85
15Gemma 4 31BGoogle85
16Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team85
17DeepSeek-R1-0528DeepSeek85
18DeepSeek-V3.2-ExpDeepSeek85
19DeepSeek-V3.2DeepSeek85
20DeepSeek-V3.2 (Thinking)DeepSeek85
21MAI-Thinking-1Microsoft85
22MiMo-V2-FlashXiaomi85
23GLM-4.5Zhipu AI85
24Kimi K2-Thinking-0905Moonshot AI85
25Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team84
26GLM-4.7Zhipu AI84
27K-EXAONE-236B-A23BLG AI Research84
28Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team84
29Nemotron 3 Super (120B A12B)NVIDIA84
30DeepSeek-V3.1DeepSeek84
31Qwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team83
32Qwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team83
33LongCat-Flash-ChatMeituan83
34Gemma 4 26B-A4BGoogle83
35LongCat-Flash-ThinkingMeituan83
36Kimi K2 0905Moonshot AI83
37Qwen3.5-9BAlibaba Cloud / Qwen Team83
38Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team82
39MiniMax M2MiniMax82
40Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team82
41Sarvam-105BSarvam AI82
42Nova 2 ProAmazon82
43GLM-4.5-AirZhipu AI81
44DeepSeek-V3 0324DeepSeek81
45Kimi K2 InstructMoonshot AI81
46Kimi K2-Instruct-0905Moonshot AI81
47MiniMax M1 80KMiniMax81
48Nova 2 LiteAmazon81
49Nova 2 OmniAmazon81
50GPT OSS 120B HighOpenAI81
51Qwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team81
52MiniMax M1 40KMiniMax81
53Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team81
54Llama 4 MaverickMeta81
55Sarvam-30BSarvam AI80
56Qwen3.5-4BAlibaba Cloud / Qwen Team79
57Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team79
58Nemotron 3 Nano (30B A3B)NVIDIA78
59LongCat-Flash-LiteMeituan78
60Mistral Small 4Mistral AI78

60 of 129 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC