MMLU

Measuring Massive Multitask Language Understanding

Models scored

100

evaluated

Modality

text

Category

finance

+6 more

Published

2020

arxiv.org

Citations

8,416

Semantic Scholar

Influential

1,616

citations

References

35

cited works

Venue

International Conference on Learning Representations

published in

Abstract

Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, et al. (+3)

We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models must possess extensive world knowledge and problem solving ability. We find that while most recent models have near random-chance accuracy, the very largest GPT-3 model improves over random chance by almost 20 percentage points on average. However, on every one of the 57 tasks, the best models still need substantial improvements before they can reach expert-level accuracy. Models also have lopsided performance and frequently do not know when they are wrong. Worse, they still have near-random accuracy on some socially important subjects such as morality and law. By comprehensively evaluating the breadth and depth of a model's academic and professional understanding, our test can be used to analyze models across many tasks and to identify important shortcomings.

financegeneralhealthcarelanguagelegalmathreasoning

Search

#ModelLabScore
01GPT-5OpenAI93
02o1OpenAI92
03o1-previewOpenAI91
04GPT-4.5OpenAI91
05Sarvam-105BSarvam AI91
06Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team91
07Claude 3.5 SonnetAnthropic90
08Claude 3.5 SonnetAnthropic90
09Kimi K2 0905Moonshot AI90
10GPT-4.1OpenAI90
11GPT OSS 120BOpenAI90
12LongCat-Flash-ChatMeituan90
13Kimi K2 InstructMoonshot AI90
14Kimi K2-Instruct-0905Moonshot AI90
15MiMo-V2.5-ProXiaomi89
16Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team89
17Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team89
18GPT-4oOpenAI89
19DeepSeek-V3DeepSeek89
20Qwen3 235B A22BAlibaba Cloud / Qwen Team88
21Kimi K2 BaseMoonshot AI88
22Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team88
23GPT-4.1 miniOpenAI88
24Grok-2xAI88
25Kimi-k1.5Moonshot AI87
26Llama 3.1 405B InstructMeta87
27o3-miniOpenAI87
28Claude 3 OpusAnthropic87
29GPT-4 TurboOpenAI87
30Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team86
31GPT-4OpenAI86
32Grok-2 minixAI86
33Llama 3.3 70B InstructMeta86
34Llama 3.2 90B InstructMeta86
35Gemini 1.5 ProGoogle86
36Nova ProAmazon86
37GPT-4oOpenAI86
38LongCat-Flash-LiteMeituan86
39Llama 4 MaverickMeta86
40GPT OSS 20BOpenAI85
41Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team85
42o1-miniOpenAI85
43Sarvam-30BSarvam AI85
44Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team85
45Phi 4Microsoft85
46Mistral Large 2Mistral AI84
47Llama 3.1 70B InstructMeta84
48Qwen2.5 32B InstructAlibaba Cloud / Qwen Team83
49Qwen2 72B InstructAlibaba Cloud / Qwen Team82
50GPT-4o miniOpenAI82
51Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team82
52Grok-1.5xAI81
53Jamba 1.5 LargeAI21 Labs81
54Mistral Small 3.1 24B BaseMistral AI81
55Mistral Small 3 24B BaseMistral AI81
56Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team81
57Mistral Small 3.1 24B InstructMistral AI81
58Nova LiteAmazon81
59Mistral Small 3.2 24B InstructMistral AI81
60DeepSeek-V2.5DeepSeek80

60 of 100 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC