ChartQA

A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Models scored

24

evaluated

Modality

multimodal

Category

multimodal

+2 more

Published

2022

arxiv.org

Citations

1,674

Semantic Scholar

Influential

256

citations

References

43

cited works

Venue

Findings

published in

Abstract

Ahmed Masry, Do Xuan Long, J. Tan, Shafiq R. Joty, et al. (+1)

Charts are very popular for analyzing data. When exploring charts, people often ask a variety of complex reasoning questions that involve several logical and arithmetic operations. They also commonly refer to visual features of a chart in their questions. However, most existing datasets do not focus on such complex reasoning questions as their questions are template-based and answers come from a fixed-vocabulary. In this work, we present a large-scale benchmark covering 9.6K human-written questions as well as 23.1K questions generated from human-written chart summaries. To address the unique challenges in our benchmark involving visual and logical reasoning over charts, we present two transformer-based models that combine visual features and the data table of the chart in a unified way to answer questions. While our models achieve the state-of-the-art results on the previous datasets as well as on our benchmark, the evaluation also reveals several challenges in answering complex reasoning questions.

Search

#ModelLabScore
01Claude 3.5 SonnetAnthropic91
02Llama 4 MaverickMeta90
03Qwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team90
04Nova ProAmazon89
05Llama 4 ScoutMeta89
06Qwen2-VL-72B-InstructAlibaba Cloud / Qwen Team88
07Pixtral LargeMistral AI88
08Mistral Small 3.2 24B InstructMistral AI87
09Qwen2.5 VL 7B InstructAlibaba Cloud / Qwen Team87
10Nova LiteAmazon87
11DeepSeek VL2DeepSeek86
12GPT-4oOpenAI86
13Llama 3.2 90B InstructMeta86
14Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team85
15DeepSeek VL2 SmallDeepSeek85
16Llama 3.2 11B InstructMeta83
17Pixtral-12BMistral AI82
18Phi-3.5-vision-instructMicrosoft82
19Phi-4-multimodal-instructMicrosoft81
20DeepSeek VL2 TinyDeepSeek81
21Gemma 3 27BGoogle78
22Grok-1.5VxAI76
23Gemma 3 12BGoogle76
24Gemma 3 4BGoogle69

24 of 24 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC