CharXiv-R

Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Models scored

47

evaluated

Modality

multimodal

Category

multimodal

+2 more

Published

2024

arxiv.org

Citations

237

Semantic Scholar

Influential

32

citations

References

0

cited works

Venue

Neural Information Processing Systems

published in

Abstract

Zirui Wang, Mengzhou Xia, Luxi He, Howard Chen, et al. (+9)

Chart understanding plays a pivotal role when applying Multimodal Large Language Models (MLLMs) to real-world tasks such as analyzing scientific papers or financial reports. However, existing datasets often focus on oversimplified and homogeneous charts with template-based questions, leading to an over-optimistic measure of progress. We demonstrate that although open-source models can appear to outperform strong proprietary models on these benchmarks, a simple stress test with slightly different charts or questions can deteriorate performance by up to 34.5%. In this work, we propose CharXiv, a comprehensive evaluation suite involving 2,323 natural, challenging, and diverse charts from arXiv papers. CharXiv includes two types of questions: 1) descriptive questions about examining basic chart elements and 2) reasoning questions that require synthesizing information across complex visual elements in the chart. To ensure quality, all charts and questions are handpicked, curated, and verified by human experts. Our results reveal a substantial, previously underestimated gap between the reasoning skills of the strongest proprietary model (i.e., GPT-4o), which achieves 47.1% accuracy, and the strongest open-source model (i.e., InternVL Chat V1.5), which achieves 29.2%. All models lag far behind human performance of 80.5%, underscoring weaknesses in the chart understanding capabilities of existing MLLMs. We hope CharXiv facilitates future research on MLLM chart understanding by providing a more realistic and faithful measure of progress. Project page and leaderboard: https://charxiv.github.io/

multimodalreasoningvision

Search

#ModelLabScore
01Claude Mythos PreviewAnthropic93
02Kimi K3Moonshot AI91
03Claude Opus 4.7Anthropic91
04Claude Opus 4.8Anthropic90
05Gemini 3.6 FlashGoogle89
06Muse Spark 1.1Meta88
07Claude Sonnet 5Anthropic88
08Kimi K2.6Moonshot AI87
09Muse SparkMeta86
10Seed 2.1 ProByteDance86
11Qwen3.7-PlusAlibaba Cloud / Qwen Team86
12Gemini 3.5 FlashGoogle84
13Seed 2.1 TurboByteDance84
14GPT-5.2OpenAI82
15GPT-5.5 InstantOpenAI82
16Qwen3.6 PlusAlibaba Cloud / Qwen Team82
17Gemini 3 ProGoogle81
18GPT-5OpenAI81
19MiMo-V2.5Xiaomi81
20Gemini 3 FlashGoogle80
21Qwen3.5-27BAlibaba Cloud / Qwen Team80
22o3OpenAI79
23Qwen3.6-27BAlibaba Cloud / Qwen Team78
24Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team78
25Kimi K2.5Moonshot AI78
26Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team78
27Claude Opus 4.6Anthropic77
28Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team77
29Gemini 3.5 Flash-LiteGoogle77
30Gemini 3.1 Flash-LiteGoogle73
31o4-miniOpenAI72
32Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team66
33Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team65
34Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team63
35Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team62
36GPT-4oOpenAI59
37GPT-4.1 miniOpenAI57
38GPT-4.1OpenAI57
39Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team57
40GPT-4.5OpenAI55
41Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team53
42Command A+Cohere53
43Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team50
44Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team49
45Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team46
46GPT-4.1 nanoOpenAI41
47Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team40

47 of 47 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC