Category
Knowledge & Reasoning
Source
Artificial Analysis
evaluation of record
Models covered
551
in our data
Data status
Live
Top score
94.1%
best on record
Top model
GPT-5.6 Sol
OpenAI
Updated
2026-07-30
last ingest
The most challenging 198 questions from GPQA, where PhD experts achieve 65% accuracy but skilled non-experts only reach 34% despite web access.
Leaderboard
Top 20 of 551 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
The hardest 198 questions from GPQA — graduate-level biology, chemistry, and physics, selected because PhD experts in the relevant field get about 65 percent while skilled non-experts with a search engine reach only 34. Higher is better, and the frontier is now above 90 percent, comfortably past the expert baseline. That saturation is the caveat: with 198 questions, a single item is half a point, so the top ten models are separated by a few questions and ordinary run-to-run variance. It stays widely quoted because it is small and cheap to run — treat it as a floor check rather than a way to tell two frontier models apart.