ChartQA
A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Models scored
24
evaluated
Modality
multimodal
Category
multimodal
+2 more
Published
2022
arxiv.org
Citations
1,674
Semantic Scholar
Influential
256
citations
References
43
cited works
Venue
Findings
published in
Abstract
Ahmed Masry, Do Xuan Long, J. Tan, Shafiq R. Joty, et al. (+1)
Charts are very popular for analyzing data. When exploring charts, people often ask a variety of complex reasoning questions that involve several logical and arithmetic operations. They also commonly refer to visual features of a chart in their questions. However, most existing datasets do not focus on such complex reasoning questions as their questions are template-based and answers come from a fixed-vocabulary. In this work, we present a large-scale benchmark covering 9.6K human-written questions as well as 23.1K questions generated from human-written chart summaries. To address the unique challenges in our benchmark involving visual and logical reasoning over charts, we present two transformer-based models that combine visual features and the data table of the chart in a unified way to answer questions. While our models achieve the state-of-the-art results on the previous datasets as well as on our benchmark, the evaluation also reveals several challenges in answering complex reasoning questions.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Claude 3.5 Sonnet | Anthropic | 91 |
| 02 | Llama 4 Maverick | Meta | 90 |
| 03 | Qwen2.5 VL 72B Instruct | Alibaba Cloud / Qwen Team | 90 |
| 04 | Nova Pro | Amazon | 89 |
| 05 | Llama 4 Scout | Meta | 89 |
| 06 | Qwen2-VL-72B-Instruct | Alibaba Cloud / Qwen Team | 88 |
| 07 | Pixtral Large | Mistral AI | 88 |
| 08 | Mistral Small 3.2 24B Instruct | Mistral AI | 87 |
| 09 | Qwen2.5 VL 7B Instruct | Alibaba Cloud / Qwen Team | 87 |
| 10 | Nova Lite | Amazon | 87 |
| 11 | DeepSeek VL2 | DeepSeek | 86 |
| 12 | GPT-4o | OpenAI | 86 |
| 13 | Llama 3.2 90B Instruct | Meta | 86 |
| 14 | Qwen2.5-Omni-7B | Alibaba Cloud / Qwen Team | 85 |
| 15 | DeepSeek VL2 Small | DeepSeek | 85 |
| 16 | Llama 3.2 11B Instruct | Meta | 83 |
| 17 | Pixtral-12B | Mistral AI | 82 |
| 18 | Phi-3.5-vision-instruct | Microsoft | 82 |
| 19 | Phi-4-multimodal-instruct | Microsoft | 81 |
| 20 | DeepSeek VL2 Tiny | DeepSeek | 81 |
| 21 | Gemma 3 27B | 78 | |
| 22 | Grok-1.5V | xAI | 76 |
| 23 | Gemma 3 12B | 76 | |
| 24 | Gemma 3 4B | 69 |
24 of 24 models · score normalized 0–100 where available