MMMU-Pro

A More Robust Multi-discipline Multimodal Understanding Benchmark

Models scored

64

evaluated

Modality

multimodal

Category

general

+3 more

Published

2024

arxiv.org

Citations

430

Semantic Scholar

Influential

65

citations

References

67

cited works

Venue

Annual Meeting of the Association for Computational Linguistics

published in

Abstract

Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, et al. (+10)

This paper introduces MMMU-Pro, a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. MMMU-Pro rigorously assesses multimodal models' true understanding and reasoning capabilities through a three-step process based on MMMU: (1) filtering out questions answerable by text-only models, (2) augmenting candidate options, and (3) introducing a vision-only input setting where questions are embedded within images. This setting challenges AI to truly"see"and"read"simultaneously, testing a fundamental human cognitive skill of seamlessly integrating visual and textual information. Results show that model performance is substantially lower on MMMU-Pro than on MMMU, ranging from 16.8% to 26.9% across models. We explore the impact of OCR prompts and Chain of Thought (CoT) reasoning, finding that OCR prompts have minimal effect while CoT generally improves performance. MMMU-Pro provides a more rigorous evaluation tool, closely mimicking real-world scenarios and offering valuable directions for future research in multimodal AI.

generalmultimodalreasoningvision

Search

#ModelLabScore
01Gemini 3.5 FlashGoogle84
02GPT-5.5OpenAI83
03GPT-5.6 SolOpenAI83
04Seed 2.1 ProByteDance83
05Seed 2.1 TurboByteDance82
06Kimi K3Moonshot AI82
07Gemini 3 FlashGoogle81
08GPT-5.4OpenAI81
09Gemini 3 ProGoogle81
10GPT-5.6 TerraOpenAI81
11Gemini 3.1 ProGoogle81
12Muse SparkMeta80
13Kimi K2.6Moonshot AI80
14GPT-5.2OpenAI80
15Qwen3.7-PlusAlibaba Cloud / Qwen Team79
16Qwen3.6 PlusAlibaba Cloud / Qwen Team79
17Kimi K2.5Moonshot AI79
18GPT-5.6 LunaOpenAI78
19GPT-5OpenAI78
20MiniMax M3MiniMax78
21MiMo-V2.5Xiaomi78
22Claude Opus 4.6Anthropic77
23Gemma 4 31BGoogle77
24Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team77
25Gemini 3.1 Flash-LiteGoogle77
26GPT-5.4 miniOpenAI77
27o3OpenAI76
28GPT-5.5 InstantOpenAI76
29Qwen3.6-27BAlibaba Cloud / Qwen Team76
30Claude Sonnet 4.6Anthropic76
31Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team75
32Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team75
33Qwen3.5-27BAlibaba Cloud / Qwen Team75
34Gemma 4 26B-A4BGoogle74
35Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team69
36Gemma 4 12BGoogle69
37Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team68
38Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team68
39GPT-5.4 nanoOpenAI66
40Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team65
41Nova 2 ProAmazon64
42Command A+Cohere63
43Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team63
44Nova 2 LiteAmazon62
45Nova 2 OmniAmazon61
46Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team60
47Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team60
48Mistral Small 4Mistral AI60
49GPT-4oOpenAI60
50Llama 4 MaverickMeta60
51Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team57
52Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team56
53DiffusionGemma 26B-A4BGoogle54
54Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team53
55Gemma 4 E4BGoogle53
56Qwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team51
57Qwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team50
58Qwen2-VL-72B-InstructAlibaba Cloud / Qwen Team46
59Llama 3.2 90B InstructMeta45
60Gemma 4 E2BGoogle44

60 of 64 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC