DocVQA

A Dataset for VQA on Document Images

Models scored

26

evaluated

Modality

multimodal

Category

image to text

+2 more

Published

2020

arxiv.org

Citations

1,540

Semantic Scholar

Influential

234

citations

References

42

cited works

Venue

IEEE Workshop/Winter Conference on Applications of Computer Vision

published in

Abstract

Minesh Mathew, Dimosthenis Karatzas, R. Manmatha, C. V. Jawahar

We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed analysis of the dataset in comparison with similar datasets for VQA and reading comprehension is presented. We report several baseline results by adopting existing VQA and reading comprehension models. Although the existing models perform reasonably well on certain types of questions, there is large performance gap compared to human performance (94.36% accuracy). The models need to improve specifically on questions where understanding structure of the document is crucial. The dataset, code and leaderboard are available at docvqa.org

image to textmultimodalvision

Search

#ModelLabScore
01Qwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team96
02Qwen2.5 VL 7B InstructAlibaba Cloud / Qwen Team96
03Claude 3.5 SonnetAnthropic95
04Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team95
05Mistral Small 3.2 24B InstructMistral AI95
06Qwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team95
07Llama 4 MaverickMeta94
08Llama 4 ScoutMeta94
09Grok-2xAI94
10Nova ProAmazon94
11DeepSeek VL2DeepSeek93
12Pixtral LargeMistral AI93
13Phi-4-multimodal-instructMicrosoft93
14Grok-2 minixAI93
15GPT-4oOpenAI93
16Nova LiteAmazon92
17DeepSeek VL2 SmallDeepSeek92
18Pixtral-12BMistral AI91
19Llama 3.2 90B InstructMeta90
20DeepSeek VL2 TinyDeepSeek89
21Llama 3.2 11B InstructMeta88
22Gemma 3 12BGoogle87
23Gemma 3 27BGoogle87
24Grok-1.5VxAI86
25Grok-1.5xAI86
26Gemma 3 4BGoogle76

26 of 26 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC