GDPval-AA v2 Leaderboard

AgenticLive

Category

Agentic

Source

Artificial Analysis

evaluation of record

Models covered

42

in our data

Data status

Live

Top score

1,861

best on record

Top model

Claude Opus 5

Anthropic

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

GDPval-AA v2 is Artificial Analysis' evaluation framework for OpenAI's GDPval dataset. It tests AI models on real-world tasks across 44 occupations and 9 major industries. Models are given shell access and web browsing capabilities in an agentic loop via Stirrup to solve tasks, with Elo ratings derived from blind pairwise comparisons.

Leaderboard

Top 20 of 42 models we hold a score for.

1Claude Opus 5Anthropic
1,8612Claude Fable 5Anthropic
1,8153GPT-5.6 SolOpenAI
1,7484Kimi K3Moonshot AI
1,6685Claude Opus 4.8Anthropic
1,6386Claude Sonnet 5Anthropic
1,6187Claude Opus 4.6Anthropic
1,6068GPT-5.6 TerraOpenAI
1,5939GPT-5.6 LunaOpenAI
1,59210Grok 4.5xAI
1,54311Claude Opus 4.7Anthropic
1,54212MiniMax M3MiniMax
1,43113GPT-5.4OpenAI
1,42914MiMo-V2-ProXiaomi
1,42615Gemini 3.6 FlashGoogle
1,42116Claude Sonnet 4.6Anthropic
1,41717MiMo-V2-OmniXiaomi
1,41018Gemini 3.5 FlashGoogle
1,37019DeepSeek-V4-Pro-MaxDeepSeek
1,33220Qwen3.7 MaxAlibaba Cloud / Qwen Team
1,308

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Takes OpenAI’s GDPval set of real occupational tasks — 44 occupations across 9 industries — runs models as agents with a shell and a browser, and then scores them by comparing two outputs side by side without revealing which model produced which. The result is an Elo rating, not a percentage: it describes who beat whom, so only the gaps between models mean anything and there is no perfect score. Read roughly 30–40 Elo as a slight edge and a few hundred points as a decisive one. Because the ratings come from pairwise preference judgments on finished deliverables, they reward presentation alongside correctness — a well-formatted wrong answer does better here than it should.

© 2026 NYSGPT2525 LLC