Category
Agentic
Source
Artificial Analysis
evaluation of record
Models covered
42
in our data
Data status
Live
Top score
1,861
best on record
Top model
Claude Opus 5
Anthropic
Updated
2026-07-30
last ingest
GDPval-AA v2 is Artificial Analysis' evaluation framework for OpenAI's GDPval dataset. It tests AI models on real-world tasks across 44 occupations and 9 major industries. Models are given shell access and web browsing capabilities in an agentic loop via Stirrup to solve tasks, with Elo ratings derived from blind pairwise comparisons.
Leaderboard
Top 20 of 42 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Takes OpenAI’s GDPval set of real occupational tasks — 44 occupations across 9 industries — runs models as agents with a shell and a browser, and then scores them by comparing two outputs side by side without revealing which model produced which. The result is an Elo rating, not a percentage: it describes who beat whom, so only the gaps between models mean anything and there is no perfect score. Read roughly 30–40 Elo as a slight edge and a few hundred points as a decisive one. Because the ratings come from pairwise preference judgments on finished deliverables, they reward presentation alongside correctness — a well-formatted wrong answer does better here than it should.