MMLU-Pro Benchmark Leaderboard

Knowledge & ReasoningLive

Category

Knowledge & Reasoning

Source

Artificial Analysis

evaluation of record

Models covered

345

in our data

Data status

Live

Top score

89.8%

best on record

Top model

Gemini 3 Pro Preview

Google

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

An enhanced version of MMLU with 12,000 graduate-level questions across 14 subject areas, featuring ten answer options and deeper reasoning requirements.

Leaderboard

Top 20 of 345 models we hold a score for.

1Gemini 3 Pro PreviewhighGoogle
89.8%2Claude Opus 4.5ReasoningAnthropic
89.5%3Gemini 3 Pro PreviewlowGoogle
89.5%4Gemini 3 Flash PreviewReasoningGoogle
89.0%5Claude Opus 4.5Non-reasoningAnthropic
88.9%6Gemini 3 Flash PreviewNon-reasoningGoogle
88.2%7Claude 4.1 OpusReasoningAnthropic
88.0%8Claude 4.5 SonnetReasoningAnthropic
87.5%9MiniMax-M2.1MiniMax
87.5%10GPT-5.2xhighOpenAI
87.4%11Claude 4 OpusReasoningAnthropic
87.3%12GPT-5highOpenAI
87.1%13GPT-5.1highOpenAI
87.0%14GPT-5mediumOpenAI
86.7%15Grok 4SpaceXAI
86.6%16GPT-5 CodexhighOpenAI
86.5%17DeepSeek V3.2 SpecialeDeepSeek
86.3%18DeepSeek V3.2ReasoningDeepSeek
86.2%19Gemini 2.5 ProGoogle
86.2%20Claude 4 OpusNon-reasoningAnthropic
86.0%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Twelve thousand graduate-level questions across 14 subjects, rebuilt from MMLU with ten answer options instead of four and questions that need reasoning rather than recall. Higher is better; the extra options alone drop the score from random guessing from 25 percent to 10, so the same model reads much lower here than on original MMLU. It is broad and stable enough to be a decent general-knowledge yardstick, and it is now old enough that contamination — test questions turning up in training data — is a live concern rather than a theoretical one. Use it to sort models into tiers, not to separate two frontier models a point apart.

© 2026 NYSGPT2525 LLC