MathVista

Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Models scored

39

evaluated

Modality

multimodal

Category

math

+2 more

Published

2023

arxiv.org

Citations

1,605

Semantic Scholar

Influential

211

citations

References

96

cited works

Venue

International Conference on Learning Representations

published in

Abstract

Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, et al. (+6)

Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at https://mathvista.github.io/.

mathmultimodalvision

Search

#ModelLabScore
01Seed 2.1 ProByteDance91
02Seed 2.1 TurboByteDance91
03o3OpenAI87
04o4-miniOpenAI84
05Step3-VL-10BStepFun84
06Command A+Cohere81
07Kimi-k1.5Moonshot AI75
08Llama 4 MaverickMeta74
09GPT-4.1 miniOpenAI73
10GPT-4.5OpenAI72
11GPT-4.1OpenAI72
12o1OpenAI72
13QvQ-72B-PreviewAlibaba Cloud / Qwen Team71
14Llama 4 ScoutMeta71
15Pixtral LargeMistral AI69
16Grok-2xAI69
17Grok-2 minixAI68
18Gemini 1.5 ProGoogle68
19Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team68
20Claude 3.5 SonnetAnthropic68
21Mistral Small 3.2 24B InstructMistral AI67
22Gemini 1.5 FlashGoogle66
23GPT-4oOpenAI64
24DeepSeek VL2DeepSeek63
25Phi-4-multimodal-instructMicrosoft62
26GPT-4oOpenAI61
27DeepSeek VL2 SmallDeepSeek61
28Pixtral-12BMistral AI58
29Llama 3.2 90B InstructMeta57
30GPT-4o miniOpenAI57
31GPT-4.1 nanoOpenAI56
32Gemini 1.5 Flash 8BGoogle55
33DeepSeek VL2 TinyDeepSeek54
34Grok-1.5VxAI53
35Grok-1.5xAI53
36Llama 3.2 11B InstructMeta52
37Gemini 1.0 ProGoogle47
38Phi-3.5-vision-instructMicrosoft44
39GPT-3.5 TurboOpenAI0

39 of 39 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC