GPQA Diamond Benchmark Leaderboard

Knowledge & ReasoningLive

Category

Knowledge & Reasoning

Source

Artificial Analysis

evaluation of record

Models covered

551

in our data

Data status

Live

Top score

94.1%

best on record

Top model

GPT-5.6 Sol

OpenAI

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

The most challenging 198 questions from GPQA, where PhD experts achieve 65% accuracy but skilled non-experts only reach 34% despite web access.

Leaderboard

Top 20 of 551 models we hold a score for.

1GPT-5.6 SolmaxOpenAI
94.1%2Gemini 3.1 Pro PreviewGoogle
94.1%3Claude Opus 5Adaptive Reasoning, High EffortAnthropic
93.7%4Claude Opus 5Adaptive Reasoning, Xhigh EffortAnthropic
93.7%5GPT-5.5xhighOpenAI
93.5%6Kimi K3Kimi
93.5%7Claude Opus 5Adaptive Reasoning, Max EffortAnthropic
93.2%8GPT-5.5highOpenAI
93.2%9GPT-5.6 SolxhighOpenAI
93.1%10Grok 4.5highSpaceXAI
93.1%11MiniMax-M3MiniMax
92.9%12GPT-5.6 SolhighOpenAI
92.8%13Gemini 3.6 FlashhighGoogle
92.8%14Claude Fable 5Adaptive Reasoning, Max Effort, Opus 4.8 FallbackAnthropic
92.6%15GPT-5.5mediumOpenAI
92.6%16GPT-5.6 SolmediumOpenAI
92.6%17GPT-5.6 TerramaxOpenAI
92.5%18Qwen3.7 MaxAlibaba
92.3%19Gemini 3.5 FlashhighGoogle
92.2%20Gemini 3.5 FlashmediumGoogle
92.1%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

The hardest 198 questions from GPQA — graduate-level biology, chemistry, and physics, selected because PhD experts in the relevant field get about 65 percent while skilled non-experts with a search engine reach only 34. Higher is better, and the frontier is now above 90 percent, comfortably past the expert baseline. That saturation is the caveat: with 198 questions, a single item is half a point, so the top ten models are separated by a few questions and ordinary run-to-run variance. It stays widely quoted because it is small and cheap to run — treat it as a floor check rather than a way to tell two frontier models apart.

© 2026 NYSGPT2525 LLC