GPQA

A Graduate-Level Google-Proof Q&A Benchmark

Models scored

232

evaluated

Modality

text

Category

biology

+4 more

Published

2023

arxiv.org

Citations

3,038

Semantic Scholar

Influential

540

citations

References

56

cited works

Venue

arXiv.org

published in

Abstract

David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, et al. (+4)

We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect), while highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web (i.e., the questions are"Google-proof"). The questions are also difficult for state-of-the-art AI systems, with our strongest GPT-4 based baseline achieving 39% accuracy. If we are to use future AI systems to help us answer very hard questions, for example, when developing new scientific knowledge, we need to develop scalable oversight methods that enable humans to supervise their outputs, which may be difficult even if the supervisors are themselves skilled and knowledgeable. The difficulty of GPQA both for skilled non-experts and frontier AI systems should enable realistic scalable oversight experiments, which we hope can help devise ways for human experts to reliably get truthful information from AI systems that surpass human capabilities.

biologychemistrygeneralphysicsreasoning

Search

#ModelLabScore
01Claude Mythos PreviewAnthropic95
02GPT-5.6 SolOpenAI95
03Gemini 3.1 ProGoogle94
04Claude Opus 4.7Anthropic94
05GPT-5.5OpenAI94
06Claude Opus 4.8Anthropic94
07Kimi K3Moonshot AI94
08GPT-5.2 ProOpenAI93
09Grok 4.5xAI93
10GPT-5.6 TerraOpenAI93
11GPT-5.4OpenAI93
12Qwen3.7 MaxAlibaba Cloud / Qwen Team92
13GPT-5.2OpenAI92
14GPT-5.6 LunaOpenAI92
15Gemini 3 ProGoogle92
16Claude Opus 4.6Anthropic91
17GLM-5.2Zhipu AI91
18Kimi K2.6Moonshot AI91
19Qwen3.6 PlusAlibaba Cloud / Qwen Team90
20Gemini 3 FlashGoogle90
21Hy3Tencent90
22Qwen3.7-PlusAlibaba Cloud / Qwen Team90
23DeepSeek-V4-Pro-MaxDeepSeek90
24Claude Sonnet 4.6Anthropic90
25Muse SparkMeta90
26Seed 2.0 ProByteDance89
27Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team88
28Grok-4 HeavyxAI88
29GPT-5.1 InstantOpenAI88
30GPT-5.1 ThinkingOpenAI88
31GPT-5.1 HighOpenAI88
32GPT-5.1OpenAI88
33GPT-5 MediumOpenAI88
34DeepSeek-V4-Flash-MaxDeepSeek88
35GPT-5.4 miniOpenAI88
36Qwen3.6-27BAlibaba Cloud / Qwen Team88
37Kimi K2.5Moonshot AI88
38Grok-4xAI88
39GPT-5 HighOpenAI87
40Claude Opus 4.5Anthropic87
41Nemotron 3 Ultra (550B A55B)NVIDIA87
42Gemini 3.1 Flash-LiteGoogle87
43Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team87
44Gemini 2.5 Pro Preview 06-05Google86
45GLM-5.1Zhipu AI86
46Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team86
47Grok 4 FastxAI86
48GPT-5OpenAI86
49GLM-4.7Zhipu AI86
50GPT-5.5 InstantOpenAI86
51Qwen3.5-27BAlibaba Cloud / Qwen Team86
52Seed 2.0 LiteByteDance85
53ERNIE 5.0Baidu85
54Claude 3.7 SonnetAnthropic85
55Grok-3xAI85
56MAI-Code-1-FlashMicrosoft85
57Kimi K2-Thinking-0905Moonshot AI85
58Gemma 4 31BGoogle84
59Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team84
60MAI-Thinking-1Microsoft84

60 of 232 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC