DROP
A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
Models scored
30
evaluated
Modality
text
Category
math
+1 more
Published
2019
arxiv.org
Citations
1,332
Semantic Scholar
Influential
232
citations
References
54
cited works
Venue
North American Chapter of the Association for Computational Linguistics
published in
Abstract
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, et al. (+2)
Reading comprehension has recently seen rapid progress, with systems matching humans on the most popular datasets for the task. However, a large body of work has highlighted the brittleness of these systems, showing that there is much work left to be done. We introduce a new reading comprehension benchmark, DROP, which requires Discrete Reasoning Over the content of Paragraphs. In this crowdsourced, adversarially-created, 55k-question benchmark, a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over them (such as addition, counting, or sorting). These operations require a much more comprehensive understanding of the content of paragraphs, as they remove the paraphrase-and-entity-typing shortcuts available in prior datasets. We apply state-of-the-art methods from both the reading comprehension and semantic parsing literatures on this dataset and show that the best systems only achieve 38.4% F1 on our generalized accuracy metric, while expert human performance is 96%. We additionally present a new model that combines reading comprehension methods with simple numerical reasoning to achieve 51% F1.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | DeepSeek-V3 | DeepSeek | 92 |
| 02 | Claude 3.5 Sonnet | Anthropic | 87 |
| 03 | Claude 3.5 Sonnet | Anthropic | 87 |
| 04 | MiMo-V2.5-Pro | Xiaomi | 86 |
| 05 | GPT-4 Turbo | OpenAI | 86 |
| 06 | Nova Pro | Amazon | 85 |
| 07 | Llama 3.1 405B Instruct | Meta | 85 |
| 08 | GPT-4o | OpenAI | 83 |
| 09 | Claude 3.5 Haiku | Anthropic | 83 |
| 10 | Claude 3 Opus | Anthropic | 83 |
| 11 | GPT-4 | OpenAI | 81 |
| 12 | Nova Lite | Amazon | 80 |
| 13 | GPT-4o mini | OpenAI | 80 |
| 14 | Llama 3.1 70B Instruct | Meta | 80 |
| 15 | Nova Micro | Amazon | 79 |
| 16 | LongCat-Flash-Chat | Meituan | 79 |
| 17 | Claude 3 Sonnet | Anthropic | 79 |
| 18 | Claude 3 Haiku | Anthropic | 78 |
| 19 | Phi 4 | Microsoft | 76 |
| 20 | Gemini 1.5 Pro | 75 | |
| 21 | GPT-3.5 Turbo | OpenAI | 70 |
| 22 | Gemma 3n E4B Instructed LiteRT Preview | 61 | |
| 23 | Gemma 3n E4B | 61 | |
| 24 | Llama 3.1 8B Instruct | Meta | 60 |
| 25 | Granite 3.3 8B Instruct | IBM | 59 |
| 26 | Gemma 3n E2B | 54 | |
| 27 | Gemma 3n E2B Instructed LiteRT (Preview) | 54 | |
| 28 | IBM Granite 4.0 Tiny Preview | IBM | 46 |
| 29 | Granite 3.3 8B Base | IBM | 36 |
| 30 | ERNIE 4.5 | Baidu | 29 |
30 of 30 models · score normalized 0–100 where available