Models scored
26
evaluated
Modality
multimodal
Category
image to text
+2 more
Published
2020
arxiv.org
Citations
1,540
Semantic Scholar
Influential
234
citations
References
42
cited works
Venue
IEEE Workshop/Winter Conference on Applications of Computer Vision
published in
Abstract
Minesh Mathew, Dimosthenis Karatzas, R. Manmatha, C. V. Jawahar
We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed analysis of the dataset in comparison with similar datasets for VQA and reading comprehension is presented. We report several baseline results by adopting existing VQA and reading comprehension models. Although the existing models perform reasonably well on certain types of questions, there is large performance gap compared to human performance (94.36% accuracy). The models need to improve specifically on questions where understanding structure of the document is crucial. The dataset, code and leaderboard are available at docvqa.org
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Qwen2.5 VL 72B Instruct | Alibaba Cloud / Qwen Team | 96 |
| 02 | Qwen2.5 VL 7B Instruct | Alibaba Cloud / Qwen Team | 96 |
| 03 | Claude 3.5 Sonnet | Anthropic | 95 |
| 04 | Qwen2.5-Omni-7B | Alibaba Cloud / Qwen Team | 95 |
| 05 | Mistral Small 3.2 24B Instruct | Mistral AI | 95 |
| 06 | Qwen2.5 VL 32B Instruct | Alibaba Cloud / Qwen Team | 95 |
| 07 | Llama 4 Maverick | Meta | 94 |
| 08 | Llama 4 Scout | Meta | 94 |
| 09 | Grok-2 | xAI | 94 |
| 10 | Nova Pro | Amazon | 94 |
| 11 | DeepSeek VL2 | DeepSeek | 93 |
| 12 | Pixtral Large | Mistral AI | 93 |
| 13 | Phi-4-multimodal-instruct | Microsoft | 93 |
| 14 | Grok-2 mini | xAI | 93 |
| 15 | GPT-4o | OpenAI | 93 |
| 16 | Nova Lite | Amazon | 92 |
| 17 | DeepSeek VL2 Small | DeepSeek | 92 |
| 18 | Pixtral-12B | Mistral AI | 91 |
| 19 | Llama 3.2 90B Instruct | Meta | 90 |
| 20 | DeepSeek VL2 Tiny | DeepSeek | 89 |
| 21 | Llama 3.2 11B Instruct | Meta | 88 |
| 22 | Gemma 3 12B | 87 | |
| 23 | Gemma 3 27B | 87 | |
| 24 | Grok-1.5V | xAI | 86 |
| 25 | Grok-1.5 | xAI | 86 |
| 26 | Gemma 3 4B | 76 |
26 of 26 models · score normalized 0–100 where available