Models scored
47
evaluated
Modality
multimodal
Category
multimodal
+2 more
Published
2024
arxiv.org
Citations
237
Semantic Scholar
Influential
32
citations
References
0
cited works
Venue
Neural Information Processing Systems
published in
Abstract
Zirui Wang, Mengzhou Xia, Luxi He, Howard Chen, et al. (+9)
Chart understanding plays a pivotal role when applying Multimodal Large Language Models (MLLMs) to real-world tasks such as analyzing scientific papers or financial reports. However, existing datasets often focus on oversimplified and homogeneous charts with template-based questions, leading to an over-optimistic measure of progress. We demonstrate that although open-source models can appear to outperform strong proprietary models on these benchmarks, a simple stress test with slightly different charts or questions can deteriorate performance by up to 34.5%. In this work, we propose CharXiv, a comprehensive evaluation suite involving 2,323 natural, challenging, and diverse charts from arXiv papers. CharXiv includes two types of questions: 1) descriptive questions about examining basic chart elements and 2) reasoning questions that require synthesizing information across complex visual elements in the chart. To ensure quality, all charts and questions are handpicked, curated, and verified by human experts. Our results reveal a substantial, previously underestimated gap between the reasoning skills of the strongest proprietary model (i.e., GPT-4o), which achieves 47.1% accuracy, and the strongest open-source model (i.e., InternVL Chat V1.5), which achieves 29.2%. All models lag far behind human performance of 80.5%, underscoring weaknesses in the chart understanding capabilities of existing MLLMs. We hope CharXiv facilitates future research on MLLM chart understanding by providing a more realistic and faithful measure of progress. Project page and leaderboard: https://charxiv.github.io/
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Claude Mythos Preview | Anthropic | 93 |
| 02 | Kimi K3 | Moonshot AI | 91 |
| 03 | Claude Opus 4.7 | Anthropic | 91 |
| 04 | Claude Opus 4.8 | Anthropic | 90 |
| 05 | Gemini 3.6 Flash | 89 | |
| 06 | Muse Spark 1.1 | Meta | 88 |
| 07 | Claude Sonnet 5 | Anthropic | 88 |
| 08 | Kimi K2.6 | Moonshot AI | 87 |
| 09 | Muse Spark | Meta | 86 |
| 10 | Seed 2.1 Pro | ByteDance | 86 |
| 11 | Qwen3.7-Plus | Alibaba Cloud / Qwen Team | 86 |
| 12 | Gemini 3.5 Flash | 84 | |
| 13 | Seed 2.1 Turbo | ByteDance | 84 |
| 14 | GPT-5.2 | OpenAI | 82 |
| 15 | GPT-5.5 Instant | OpenAI | 82 |
| 16 | Qwen3.6 Plus | Alibaba Cloud / Qwen Team | 82 |
| 17 | Gemini 3 Pro | 81 | |
| 18 | GPT-5 | OpenAI | 81 |
| 19 | MiMo-V2.5 | Xiaomi | 81 |
| 20 | Gemini 3 Flash | 80 | |
| 21 | Qwen3.5-27B | Alibaba Cloud / Qwen Team | 80 |
| 22 | o3 | OpenAI | 79 |
| 23 | Qwen3.6-27B | Alibaba Cloud / Qwen Team | 78 |
| 24 | Qwen3.6-35B-A3B | Alibaba Cloud / Qwen Team | 78 |
| 25 | Kimi K2.5 | Moonshot AI | 78 |
| 26 | Qwen3.5-35B-A3B | Alibaba Cloud / Qwen Team | 78 |
| 27 | Claude Opus 4.6 | Anthropic | 77 |
| 28 | Qwen3.5-122B-A10B | Alibaba Cloud / Qwen Team | 77 |
| 29 | Gemini 3.5 Flash-Lite | 77 | |
| 30 | Gemini 3.1 Flash-Lite | 73 | |
| 31 | o4-mini | OpenAI | 72 |
| 32 | Qwen3 VL 235B A22B Thinking | Alibaba Cloud / Qwen Team | 66 |
| 33 | Qwen3 VL 32B Thinking | Alibaba Cloud / Qwen Team | 65 |
| 34 | Qwen3 VL 32B Instruct | Alibaba Cloud / Qwen Team | 63 |
| 35 | Qwen3 VL 235B A22B Instruct | Alibaba Cloud / Qwen Team | 62 |
| 36 | GPT-4o | OpenAI | 59 |
| 37 | GPT-4.1 mini | OpenAI | 57 |
| 38 | GPT-4.1 | OpenAI | 57 |
| 39 | Qwen3 VL 30B A3B Thinking | Alibaba Cloud / Qwen Team | 57 |
| 40 | GPT-4.5 | OpenAI | 55 |
| 41 | Qwen3 VL 8B Thinking | Alibaba Cloud / Qwen Team | 53 |
| 42 | Command A+ | Cohere | 53 |
| 43 | Qwen3 VL 4B Thinking | Alibaba Cloud / Qwen Team | 50 |
| 44 | Qwen3 VL 30B A3B Instruct | Alibaba Cloud / Qwen Team | 49 |
| 45 | Qwen3 VL 8B Instruct | Alibaba Cloud / Qwen Team | 46 |
| 46 | GPT-4.1 nano | OpenAI | 41 |
| 47 | Qwen3 VL 4B Instruct | Alibaba Cloud / Qwen Team | 40 |
47 of 47 models · score normalized 0–100 where available