DROP

A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Models scored

30

evaluated

Modality

text

Category

math

+1 more

Published

2019

arxiv.org

Citations

1,332

Semantic Scholar

Influential

232

citations

References

54

cited works

Venue

North American Chapter of the Association for Computational Linguistics

published in

Abstract

Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, et al. (+2)

Reading comprehension has recently seen rapid progress, with systems matching humans on the most popular datasets for the task. However, a large body of work has highlighted the brittleness of these systems, showing that there is much work left to be done. We introduce a new reading comprehension benchmark, DROP, which requires Discrete Reasoning Over the content of Paragraphs. In this crowdsourced, adversarially-created, 55k-question benchmark, a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over them (such as addition, counting, or sorting). These operations require a much more comprehensive understanding of the content of paragraphs, as they remove the paraphrase-and-entity-typing shortcuts available in prior datasets. We apply state-of-the-art methods from both the reading comprehension and semantic parsing literatures on this dataset and show that the best systems only achieve 38.4% F1 on our generalized accuracy metric, while expert human performance is 96%. We additionally present a new model that combines reading comprehension methods with simple numerical reasoning to achieve 51% F1.

Search

#ModelLabScore
01DeepSeek-V3DeepSeek92
02Claude 3.5 SonnetAnthropic87
03Claude 3.5 SonnetAnthropic87
04MiMo-V2.5-ProXiaomi86
05GPT-4 TurboOpenAI86
06Nova ProAmazon85
07Llama 3.1 405B InstructMeta85
08GPT-4oOpenAI83
09Claude 3.5 HaikuAnthropic83
10Claude 3 OpusAnthropic83
11GPT-4OpenAI81
12Nova LiteAmazon80
13GPT-4o miniOpenAI80
14Llama 3.1 70B InstructMeta80
15Nova MicroAmazon79
16LongCat-Flash-ChatMeituan79
17Claude 3 SonnetAnthropic79
18Claude 3 HaikuAnthropic78
19Phi 4Microsoft76
20Gemini 1.5 ProGoogle75
21GPT-3.5 TurboOpenAI70
22Gemma 3n E4B Instructed LiteRT PreviewGoogle61
23Gemma 3n E4BGoogle61
24Llama 3.1 8B InstructMeta60
25Granite 3.3 8B InstructIBM59
26Gemma 3n E2BGoogle54
27Gemma 3n E2B Instructed LiteRT (Preview)Google54
28IBM Granite 4.0 Tiny PreviewIBM46
29Granite 3.3 8B BaseIBM36
30ERNIE 4.5Baidu29

30 of 30 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC