Category
Knowledge & Reasoning
Source
Artificial Analysis
evaluation of record
Models covered
345
in our data
Data status
Live
Top score
89.8%
best on record
Top model
Gemini 3 Pro Preview
Updated
2026-07-30
last ingest
An enhanced version of MMLU with 12,000 graduate-level questions across 14 subject areas, featuring ten answer options and deeper reasoning requirements.
Leaderboard
Top 20 of 345 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Twelve thousand graduate-level questions across 14 subjects, rebuilt from MMLU with ten answer options instead of four and questions that need reasoning rather than recall. Higher is better; the extra options alone drop the score from random guessing from 25 percent to 10, so the same model reads much lower here than on original MMLU. It is broad and stable enough to be a decent general-knowledge yardstick, and it is now old enough that contamination — test questions turning up in training data — is a live concern rather than a theoretical one. Use it to sort models into tiers, not to separate two frontier models a point apart.