MMMU-Pro Benchmark Leaderboard

Multilingual & MultimodalLive

Category

Multilingual & Multimodal

Source

Artificial Analysis

evaluation of record

Models covered

64

in our data

Data status

Live

Top score

83.6%

best on record

Top model

Gemini 3.5 Flash

Google

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

An enhanced MMMU benchmark that eliminates shortcuts and guessing strategies to more rigorously test multimodal models across 30 academic disciplines.

Leaderboard

Top 20 of 64 models we hold a score for.

1Gemini 3.5 FlashGoogle
83.6%2GPT-5.5OpenAI
83.2%3GPT-5.6 SolOpenAI
83.0%4Seed 2.1 ProByteDance
82.7%5Seed 2.1 TurboByteDance
82.2%6Kimi K3Moonshot AI
81.6%7Gemini 3 FlashGoogle
81.2%8GPT-5.4OpenAI
81.2%9Gemini 3 ProGoogle
81.0%10GPT-5.6 TerraOpenAI
80.7%11Gemini 3.1 ProGoogle
80.5%12Muse SparkMeta
80.4%13Kimi K2.6Moonshot AI
80.1%14GPT-5.2OpenAI
79.5%15Qwen3.7-PlusAlibaba Cloud / Qwen Team
79.0%16Qwen3.6 PlusAlibaba Cloud / Qwen Team
78.8%17Kimi K2.5Moonshot AI
78.5%18GPT-5.6 LunaOpenAI
78.4%19GPT-5OpenAI
78.4%20MiniMax M3MiniMax
78.1%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

A harder rebuild of MMMU: college-level questions across 30 disciplines that require reading an image, with the shortcuts stripped out — more answer options, and a version where the question itself is embedded in the screenshot so the model cannot skip the vision step. Higher is better; the same model typically scores 15 to 20 points below plain MMMU, which is exactly what the redesign was for. A high score here means the model genuinely read the diagram rather than pattern-matching the surrounding text. Vision results are more sensitive to image resolution and prompt format than text results, so comparisons across labs deserve a look at how each run was configured.

© 2026 NYSGPT2525 LLC