PolyMATH

A Challenging Multi-modal Mathematical Reasoning Benchmark

Models scored

23

evaluated

Modality

multimodal

Category

math

+4 more

Published

2024

arxiv.org

Citations

22

Semantic Scholar

Influential

2

citations

References

134

cited works

Venue

arXiv.org

published in

Abstract

Himanshu Gupta, Shreyas Verma, Ujjwala Anantheswaran, Kevin Scaria, et al. (+3)

Multi-modal Large Language Models (MLLMs) exhibit impressive problem-solving abilities in various domains, but their visual comprehension and abstract reasoning skills remain under-evaluated. To this end, we present PolyMATH, a challenging benchmark aimed at evaluating the general cognitive reasoning abilities of MLLMs. PolyMATH comprises 5,000 manually collected high-quality images of cognitive textual and visual challenges across 10 distinct categories, including pattern recognition, spatial reasoning, and relative reasoning. We conducted a comprehensive, and quantitative evaluation of 15 MLLMs using four diverse prompting strategies, including Chain-of-Thought and Step-Back. The best scores achieved on PolyMATH are ~41%, ~36%, and ~27%, obtained by Claude-3.5 Sonnet, GPT-4o and Gemini-1.5 Pro respectively - highlighting the logical and visual complexity of these questions. A further fine-grained error analysis reveals that these models struggle to understand spatial relations and perform drawn-out, high-level reasoning. This is further strengthened by our ablation study estimating MLLM performance when given textual descriptions in place of diagrams. As evidenced by ~4% improvement over textual descriptions as opposed to actual images, we discover that models do not truly comprehend visual diagrams and the spatial information therein, and are thus prone to logical errors. Finally, we evaluate the OpenAI o1 models and find that their performance only matches the human baseline, highlighting the difficulty of the benchmark. The results on PolyMATH highlight the room for improvement in multi-modal reasoning and provide unique insights to guide the development of future MLLMs.

mathmultimodalreasoningspatial reasoningvision

Search

#ModelLabScore
01Qwen3.7 MaxAlibaba Cloud / Qwen Team87
02Qwen3.7-PlusAlibaba Cloud / Qwen Team84
03Qwen3.6 PlusAlibaba Cloud / Qwen Team77
04Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team73
05Qwen3.5-27BAlibaba Cloud / Qwen Team71
06Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team69
07Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team64
08Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team60
09Qwen3.5-9BAlibaba Cloud / Qwen Team57
10Qwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team56
11Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team52
12Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team52
13Qwen3.5-4BAlibaba Cloud / Qwen Team51
14Qwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team50
15Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team48
16Qwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team46
17Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team45
18Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team44
19Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team41
20Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team30
21Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team29
22Qwen3.5-2BAlibaba Cloud / Qwen Team26
23Qwen3.5-0.8BAlibaba Cloud / Qwen Team8

23 of 23 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC