MathVista
Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Models scored
39
evaluated
Modality
multimodal
Category
math
+2 more
Published
2023
arxiv.org
Citations
1,605
Semantic Scholar
Influential
211
citations
References
96
cited works
Venue
International Conference on Learning Representations
published in
Abstract
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, et al. (+6)
Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at https://mathvista.github.io/.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Seed 2.1 Pro | ByteDance | 91 |
| 02 | Seed 2.1 Turbo | ByteDance | 91 |
| 03 | o3 | OpenAI | 87 |
| 04 | o4-mini | OpenAI | 84 |
| 05 | Step3-VL-10B | StepFun | 84 |
| 06 | Command A+ | Cohere | 81 |
| 07 | Kimi-k1.5 | Moonshot AI | 75 |
| 08 | Llama 4 Maverick | Meta | 74 |
| 09 | GPT-4.1 mini | OpenAI | 73 |
| 10 | GPT-4.5 | OpenAI | 72 |
| 11 | GPT-4.1 | OpenAI | 72 |
| 12 | o1 | OpenAI | 72 |
| 13 | QvQ-72B-Preview | Alibaba Cloud / Qwen Team | 71 |
| 14 | Llama 4 Scout | Meta | 71 |
| 15 | Pixtral Large | Mistral AI | 69 |
| 16 | Grok-2 | xAI | 69 |
| 17 | Grok-2 mini | xAI | 68 |
| 18 | Gemini 1.5 Pro | 68 | |
| 19 | Qwen2.5-Omni-7B | Alibaba Cloud / Qwen Team | 68 |
| 20 | Claude 3.5 Sonnet | Anthropic | 68 |
| 21 | Mistral Small 3.2 24B Instruct | Mistral AI | 67 |
| 22 | Gemini 1.5 Flash | 66 | |
| 23 | GPT-4o | OpenAI | 64 |
| 24 | DeepSeek VL2 | DeepSeek | 63 |
| 25 | Phi-4-multimodal-instruct | Microsoft | 62 |
| 26 | GPT-4o | OpenAI | 61 |
| 27 | DeepSeek VL2 Small | DeepSeek | 61 |
| 28 | Pixtral-12B | Mistral AI | 58 |
| 29 | Llama 3.2 90B Instruct | Meta | 57 |
| 30 | GPT-4o mini | OpenAI | 57 |
| 31 | GPT-4.1 nano | OpenAI | 56 |
| 32 | Gemini 1.5 Flash 8B | 55 | |
| 33 | DeepSeek VL2 Tiny | DeepSeek | 54 |
| 34 | Grok-1.5V | xAI | 53 |
| 35 | Grok-1.5 | xAI | 53 |
| 36 | Llama 3.2 11B Instruct | Meta | 52 |
| 37 | Gemini 1.0 Pro | 47 | |
| 38 | Phi-3.5-vision-instruct | Microsoft | 44 |
| 39 | GPT-3.5 Turbo | OpenAI | 0 |
39 of 39 models · score normalized 0–100 where available