GPQA: A Graduate-Level Google-Proof Q&A Benchmark

We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect), while highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web (i.e., the questions are"Google-proof"). The questions are also difficult for state-of-the-art AI systems, with our strongest GPT-4 based baseline achieving 39% accuracy. If we are to use future AI systems to help us answer very hard questions, for example, when developing new scientific knowledge, we need to develop scalable oversight methods that enable humans to supervise their outputs, which may be difficult even if the supervisors are themselves skilled and knowledgeable. The difficulty of GPQA both for skilled non-experts and frontier AI systems should enable realistic scalable oversight experiments, which we hope can help devise ways for human experts to reliably get truthful information from AI systems that surpass human capabilities.

Paper

References (56)

Scroll for more · 38 remaining

Similar papers

Reviewer iAEH8/10 · confidence 4/52024-04-30

Summary

This paper presents GPQA, a new dataset of challenging multi-choice question-answer pairs. GPQA was created with the help of human experts to be challenging to even experts in the field, as a means of assessing LLM performance on frontier knowledge. The authors selected as annotators people having or pursuing a PhD in the subject matter for each question (physics, chemistry, biology). Each question was created by one expert, double-checked by another expert, and then revised by the first expert. The authors then test non-experts (annotators who worked on other fields) on each question. The dataset contains several parts. The "diamond split" is a subset of 198 questions where 2/2 experts agree on the answer, and 2/3 non-experts get the answer *wrong*. The "main set" is an extension to 448 questions; these may included questions that required revision, and may include questions where only 1/3 non-expert got the answer wrong. Finally, the "extended set" includes all 546 questions created. The authors recommend the main set for experiments. For the main set, experts achieve 71.9% accuracy, while non-experts (with access to Google) achieve 30.4% accuracy. Language models tend to lie in-between, with the best-performing being Claude-3 Opus at 52.7%.

Rating

8

Confidence

4

Ethics flag

1

Reasons to accept

The dataset provides an interesting testbed for question answering performance on highly difficult problems. With each question requiring two expert annotators, substantial resources went into the creation of this dataset; it will provide a good evaluation benchmark for the community. The finding that LLMs perform better than non-experts, even when equipped with Google, is a good demonstration of their potential.

Reasons to reject

Despite the effort that went into creating this dataset, it is a small resource. With only 448 examples in the main set, it may be difficult to e.g. estimate 95% confidence intervals for model performance on the dataset.

Questions to authors

I am a little confused about the framing in the introduction. You argue that this dataset is constructed to evaluated models "on questions where we cannot produce or verify the truth on our own". Yet, we (taking the word to mean humanity) *can* answer all the questions on this dataset -- experts have this ability. Why would you expect evaluation techniques for frontier knowledge known by some humans to generalise to frontier knowledge known by no humans? If you do not have this expectation, how does this paper connect to scalable oversight?

Reviewer TwFw8/10 · confidence 4/52024-05-11

Summary

This paper presents GPQA, a highly challenging evaluation dataset of 448 multiple-choice questions in the fields of biology, physics, and chemistry. The questions are designed to be "Google-proof", i.e., with 34% accuracy by highly skilled non-experts spending over 30 minutes on average with unrestricted access to the web. The dataset aims to aid the development of scalable oversight methods for supervising AI systems that surpass human capabilities. The proposed dataset construction pipeline involving experts and non-experts ensures the difficulty and objectivity of the questions in GPQA. A high-quality subset of the dataset ,GPQA Diamond, is also proposed, where both experts answer correctly and majority of non-experts answer incorrectly. The difficulty and objectivity have also been validated by humans in the follow-up analysis. The accuracies of several LLMs with self-consistency are also reported.

Rating

8

Confidence

4

Ethics flag

1

Reasons to accept

* The proposed evaluation dataset contains difficult and objective questions. The dataset construction process and follow-up analysis ensure the quality of the dataset. * The experiments show that the dataset is difficult for even strong LLMs such as GPT-4. The evaluation results of humans and LLMs should inspire future work on "scalable oversight".

Reasons to reject

* The size of the dataset, 448, is relatively small. Such small evaluation datasets may not be adequate to ensure the statistical significance of model performance. * The focus of the dataset is limited to biology, physics, and chemistry, which is narrower compared to broad domains in practical applications of LLMs. * This paper seems to lack reference and comparison to existing evaluation datasets, which makes the contribution of this paper unclear.

Reviewer TwFw2024-06-04

Comment by reviewer

Thank you for your hard work on the paper, and the rebuttal! Comments by the authors addressed my 3rd concern, and I understood that the authors mentioned the small size in Limitation section. So I'll raise my score.

Reviewer D4AP8/10 · confidence 4/52024-05-11

Summary

The paper introduces a challenging dataset designed to test the abilities of both human experts and AI systems in solving difficult multiple-choice questions across three scientific domains: biology, physics, and chemistry. The dataset consists of 448 questions, which were found to be extremely challenging for experts and non-experts alike, designed to be inaccessible through simple Google searches, thus the term "Google-proof". The study highlights significant differences in performance between human experts, non-expert validators, and various AI models, with experts achieving a 65% accuracy, non-experts 34%, and AI models like Claude 3 Opus around 60%.

Rating

8

Confidence

4

Ethics flag

2

Reasons to accept

- Innovative Dataset: The GPQA dataset fills a critical gap in existing benchmarks by focusing on extremely hard questions that require deep domain expertise and are resistant to simple internet searches. - Rigorous Validation Process: The paper details a thorough validation process for question objectivity and difficulty, involving multiple rounds of expert and non-expert reviews, ensuring the reliability and challenge of the dataset.

Reasons to reject

- Limited Diversity of Questions: The dataset, while challenging, consists of only 448 questions due to the high costs and complexity of question generation and validation. This small size might limit the statistical power and generalizability of the findings.

Reviewer 1Knd9/10 · confidence 4/52024-05-12

Summary

Contribution: * 448 multiple choice questions in biology, physics, chemistry, written by experts * PhD’s in the corresponding domain only get 65% of the questions right, 34% for skilled non experts with Google access * Gpt-4 achieves 39% accuracy, Claude 3 Opua 60% * GPQA may enable “scalable oversight” because of its high difficulty GPQA seems to be a very useful benchmark that will benefit the community.

Rating

9

Confidence

4

Ethics flag

1

Reasons to accept

Looking at the examples in Table 1, the examples in this task seem to be of great quality and incredibly difficult for non experts. I'm very impressed with the quality of the data and I'm surprised that the non expert human performance is so low. The incentive structure for the data creators is very interesting. It's very useful to know exactly how much the workers were paid and how their compensation was tied to performance.

Reasons to reject

The fact that the questions in GPQA are multiple choice could be a big limitation of this dataset, mainly because it might make it non realistic. By this I mean that real world researchers will likely never face tricky multiple choice questions in their work. Writing useful questions is much more likely to be a useful research task. There is additionally a real risk that expert question writers found ways to game the rules of the data collection to maximize their pay, making the dataset even less realistic. This leads me to wonder if the reason that in domain PhDs can't solve this task with a high degree of accuracy may actually mean that the questions are needlessly obscure rather than difficult in a useful way.

Questions to authors

What accuracy would in domain PhDs have if they had 30 minutes and access to the internet? It would be interesting to see some ablations for the data collection choices. For example, how much do performance based financial incentives improve data quality? It would be really interesting to measure how useful LLMs can be to experts or non experts that are trying to solve difficult questions like the ones in GPQA. Can the LLM output accurate and comprehensive rationales that are useful to a human's understanding?

Reviewer 1Knd2024-06-04

Thanks for the responses

Authorsrebuttal2024-06-06

Let us know if you have any more questions about the motivation and context on scalable oversight—happy to discuss more!

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC