Category
Multilingual & Multimodal
Source
Artificial Analysis
evaluation of record
Models covered
64
in our data
Data status
Live
Top score
83.6%
best on record
Top model
Gemini 3.5 Flash
Updated
2026-07-30
last ingest
An enhanced MMMU benchmark that eliminates shortcuts and guessing strategies to more rigorously test multimodal models across 30 academic disciplines.
Leaderboard
Top 20 of 64 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
A harder rebuild of MMMU: college-level questions across 30 disciplines that require reading an image, with the shortcuts stripped out — more answer options, and a version where the question itself is embedded in the screenshot so the model cannot skip the vision step. Higher is better; the same model typically scores 15 to 20 points below plain MMMU, which is exactly what the redesign was for. A high score here means the model genuinely read the diagram rather than pattern-matching the surrounding text. Vision results are more sensitive to image resolution and prompt format than text results, so comparisons across labs deserve a look at how each run was configured.