MMMU

A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Models scored

63

evaluated

Modality

multimodal

Category

general

+4 more

Published

2023

arxiv.org

Citations

2,204

Semantic Scholar

Influential

296

citations

References

84

cited works

Venue

Computer Vision and Pattern Recognition

published in

Abstract

Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, et al. (+18)

We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and text-books, covering six core disciplines: Art & Design, Busi-ness, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly het-erogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. The evaluation of 28 open-source LMMs as well as the propri-etary GPT-4V(ision) and Gemini highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V and Gemini Ultra only achieve accuracies of 56% and 59% respectively, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence.

generalhealthcaremultimodalreasoningvision

Search

#ModelLabScore
01Qwen3.6 PlusAlibaba Cloud / Qwen Team86
02GPT-5.1OpenAI85
03GPT-5.1 InstantOpenAI85
04GPT-5.1 ThinkingOpenAI85
05GPT-5OpenAI84
06Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team84
07o3OpenAI83
08Qwen3.6-27BAlibaba Cloud / Qwen Team83
09Qwen3.5-27BAlibaba Cloud / Qwen Team82
10Gemini 2.5 Pro Preview 06-05Google82
11Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team82
12o4-miniOpenAI82
13Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team81
14Gemini 2.5 FlashGoogle80
15Gemini 2.5 ProGoogle80
16Step3-VL-10BStepFun78
17Grok-3xAI78
18o1OpenAI78
19Gemini 2.0 Flash ThinkingGoogle75
20GPT-4.5OpenAI75
21Command A+Cohere75
22Claude 3.7 SonnetAnthropic75
23GPT-4.1OpenAI75
24Claude Sonnet 4Anthropic74
25Llama 4 MaverickMeta73
26Gemini 2.5 Flash-LiteGoogle73
27GPT-4.1 miniOpenAI73
28GPT-4oOpenAI72
29Gemini 2.0 FlashGoogle71
30QvQ-72B-PreviewAlibaba Cloud / Qwen Team70
31Qwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team70
32Qwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team70
33Kimi-k1.5Moonshot AI70
34Llama 4 ScoutMeta69
35Claude 3.5 SonnetAnthropic68
36Gemini 2.0 Flash-LiteGoogle68
37Grok-2xAI66
38Gemini 1.5 ProGoogle66
39Pixtral LargeMistral AI64
40Grok-2 minixAI63
41Mistral Small 3.2 24B InstructMistral AI63
42Gemini 1.5 FlashGoogle62
43Nova ProAmazon62
44Llama 3.2 90B InstructMeta60
45GPT-4o miniOpenAI59
46Mistral Small 3.1 24B InstructMistral AI59
47Mistral Small 3.1 24B BaseMistral AI59
48Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team59
49Qwen2.5 VL 7B InstructAlibaba Cloud / Qwen Team59
50Nova LiteAmazon56
51GPT-4.1 nanoOpenAI55
52Phi-4-multimodal-instructMicrosoft55
53Gemini 1.5 Flash 8BGoogle54
54Grok-1.5xAI54
55Grok-1.5VxAI54
56Pixtral-12BMistral AI53
57DeepSeek VL2DeepSeek51
58Llama 3.2 11B InstructMeta51
59DeepSeek VL2 SmallDeepSeek48
60Gemini 1.0 ProGoogle48

60 of 63 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC