MathVista-Mini

Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Models scored

23

evaluated

Modality

multimodal

Category

math

+2 more

Published

2023

arxiv.org

Citations

1,605

Semantic Scholar

Influential

211

citations

References

96

cited works

Venue

International Conference on Learning Representations

published in

Abstract

Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, et al. (+6)

Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at https://mathvista.github.io/.

mathmultimodalvision

Search

#ModelLabScore
01Kimi K2.5Moonshot AI90
02Qwen3.5-27BAlibaba Cloud / Qwen Team88
03Qwen3.6-27BAlibaba Cloud / Qwen Team87
04Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team87
05Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team86
06Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team86
07Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team86
08Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team86
09Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team85
10Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team84
11Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team82
12Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team81
13Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team80
14Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team80
15Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team77
16Qwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team75
17Qwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team75
18Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team74
19Qwen2-VL-72B-InstructAlibaba Cloud / Qwen Team71
20Qwen2.5 VL 7B InstructAlibaba Cloud / Qwen Team68
21Gemma 3 27BGoogle68
22Gemma 3 12BGoogle63
23Gemma 3 4BGoogle50

23 of 23 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC