AI2D

A Diagram is Worth a Dozen Images

Models scored

32

evaluated

Modality

multimodal

Category

multimodal

+2 more

Published

2016

arxiv.org

Citations

1,030

Semantic Scholar

Influential

132

citations

References

59

cited works

Venue

European Conference on Computer Vision

published in

Abstract

Aniruddha Kembhavi, M. Salvato, Eric Kolve, Minjoon Seo, et al. (+2)

Diagrams are common tools for representing complex concepts, relationships and events, often when it would be difficult to portray the same information with natural images. Understanding natural images has been extensively studied in computer vision, while diagram understanding has received little attention. In this paper, we study the problem of diagram interpretation, the challenging task of identifying the structure of a diagram and the semantics of its constituents and their relationships. We introduce Diagram Parse Graphs (DPG) as our representation to model the structure of diagrams. We define syntactic parsing of diagrams as learning to infer DPGs for diagrams and study semantic interpretation and reasoning of diagrams in the context of diagram question answering. We devise an LSTM-based method for syntactic parsing of diagrams and introduce a DPG-based attention model for diagram question answering. We compile a new dataset of diagrams with exhaustive annotations of constituents and relationships for about 5,000 diagrams and 15,000 questions and answers. Our results show the significance of our models for syntactic parsing and question answering in diagrams using DPGs.

Search

#ModelLabScore
01Claude 3.5 SonnetAnthropic95
02Qwen3.6 PlusAlibaba Cloud / Qwen Team94
03GPT-4oOpenAI94
04Pixtral LargeMistral AI94
05Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team93
06Mistral Small 3.2 24B InstructMistral AI93
07Qwen3.5-27BAlibaba Cloud / Qwen Team93
08Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team93
09Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team93
10Llama 3.2 90B InstructMeta92
11Llama 3.2 11B InstructMeta91
12Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team90
13Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team90
14Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team89
15Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team89
16Qwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team88
17Grok-1.5VxAI88
18Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team87
19Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team86
20Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team85
21Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team85
22Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team85
23Gemma 3 27BGoogle85
24Gemma 3 12BGoogle84
25Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team84
26Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team83
27Phi-4-multimodal-instructMicrosoft82
28DeepSeek VL2DeepSeek81
29DeepSeek VL2 SmallDeepSeek80
30Phi-3.5-vision-instructMicrosoft78
31Gemma 3 4BGoogle75
32DeepSeek VL2 TinyDeepSeek72

32 of 32 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC