MMLU-ProX

A More Robust and Challenging Multi-Task Language Understanding Benchmark

Models scored

32

evaluated

Modality

text

Category

finance

+6 more

Published

2024

arxiv.org

Citations

1,729

Semantic Scholar

Influential

197

citations

References

56

cited works

Venue

Neural Information Processing Systems

published in

Abstract

Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, et al. (+13)

In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning across diverse domains. However, as models continue to improve, their performance on these benchmarks has begun to plateau, making it increasingly difficult to discern differences in model capabilities. This paper introduces MMLU-Pro, an enhanced dataset designed to extend the mostly knowledge-driven MMLU benchmark by integrating more challenging, reasoning-focused questions and expanding the choice set from four to ten options. Additionally, MMLU-Pro eliminates the trivial and noisy questions in MMLU. Our experimental results show that MMLU-Pro not only raises the challenge, causing a significant drop in accuracy by 16% to 33% compared to MMLU but also demonstrates greater stability under varying prompts. With 24 different prompt styles tested, the sensitivity of model scores to prompt variations decreased from 4-5% in MMLU to just 2% in MMLU-Pro. Additionally, we found that models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering, which is in stark contrast to the findings on the original MMLU, indicating that MMLU-Pro includes more complex reasoning questions. Our assessments confirm that MMLU-Pro is a more discriminative benchmark to better track progress in the field.

financegeneralhealthcarelanguagelegalmathreasoning

Search

#ModelLabScore
01Qwen3.7 MaxAlibaba Cloud / Qwen Team87
02Qwen3.7-PlusAlibaba Cloud / Qwen Team85
03Qwen3.6 PlusAlibaba Cloud / Qwen Team85
04Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team85
05Nemotron 3 Ultra (550B A55B)NVIDIA83
06Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team82
07Qwen3.5-27BAlibaba Cloud / Qwen Team82
08Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team81
09Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team81
10Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team81
11Qwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team79
12Nemotron 3 Super (120B A12B)NVIDIA79
13Qwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team79
14Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team78
15Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team77
16Qwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team77
17Qwen3.5-9BAlibaba Cloud / Qwen Team76
18Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team76
19Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team73
20Qwen3.5-4BAlibaba Cloud / Qwen Team72
21Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team71
22Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team71
23Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team65
24Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team65
25Nemotron 3 Nano (30B A3B)NVIDIA60
26Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team59
27Qwen3.5-2BAlibaba Cloud / Qwen Team52
28Qwen3.5-0.8BAlibaba Cloud / Qwen Team35
29Gemma 3n E4B Instructed LiteRT PreviewGoogle20
30Gemma 3n E4B InstructedGoogle20
31Gemma 3n E2B Instructed LiteRT (Preview)Google8
32Gemma 3n E2B InstructedGoogle8

32 of 32 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC