Models scored
34
evaluated
Modality
text
Category
chemistry
+8 more
Published
2025
arxiv.org
Citations
0
Semantic Scholar
Influential
0
citations
References
0
cited works
Venue
—
published in
Abstract
M-A-P Team, Xinrun Du, Yifan Yao, Kaijing Ma, et al. (+93)
Large language models (LLMs) have demonstrated remarkable proficiency in mainstream academic disciplines such as mathematics, physics, and computer science. However, human knowledge encompasses over 200 specialized disciplines, far exceeding the scope of existing benchmarks. The capabilities of LLMs in many of these specialized fields-particularly in light industry, agriculture, and service-oriented disciplines-remain inadequately evaluated. To address this gap, we present SuperGPQA, a comprehensive benchmark that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines. Our benchmark employs a novel Human-LLM collaborative filtering mechanism to eliminate trivial or ambiguous questions through iterative refinement based on both LLM responses and expert feedback. Our experimental results reveal significant room for improvement in the performance of current state-of-the-art LLMs across diverse knowledge domains (e.g., the reasoning-focused model DeepSeek-R1 achieved the highest accuracy of 61.82% on SuperGPQA), highlighting the considerable gap between current model capabilities and artificial general intelligence. Additionally, we present comprehensive insights from our management of a large-scale annotation process, involving over 80 expert annotators and an interactive Human-LLM collaborative system, offering valuable methodological guidance for future research initiatives of comparable scope.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Qwen3.7 Max | Alibaba Cloud / Qwen Team | 74 |
| 02 | Qwen3.6 Plus | Alibaba Cloud / Qwen Team | 72 |
| 03 | Qwen3.7-Plus | Alibaba Cloud / Qwen Team | 71 |
| 04 | Seed 2.1 Pro | ByteDance | 71 |
| 05 | Qwen3.5-397B-A17B | Alibaba Cloud / Qwen Team | 70 |
| 06 | Seed 2.1 Turbo | ByteDance | 67 |
| 07 | Qwen3.5-122B-A10B | Alibaba Cloud / Qwen Team | 67 |
| 08 | Qwen3.6-27B | Alibaba Cloud / Qwen Team | 66 |
| 09 | Qwen3.5-27B | Alibaba Cloud / Qwen Team | 66 |
| 10 | Qwen3 Max | Alibaba Cloud / Qwen Team | 65 |
| 11 | Qwen3-235B-A22B-Thinking-2507 | Alibaba Cloud / Qwen Team | 65 |
| 12 | Qwen3.6-35B-A3B | Alibaba Cloud / Qwen Team | 65 |
| 13 | Qwen3 VL 235B A22B Thinking | Alibaba Cloud / Qwen Team | 64 |
| 14 | Qwen3.5-35B-A3B | Alibaba Cloud / Qwen Team | 63 |
| 15 | Qwen3-235B-A22B-Instruct-2507 | Alibaba Cloud / Qwen Team | 63 |
| 16 | Qwen3-Next-80B-A3B-Thinking | Alibaba Cloud / Qwen Team | 61 |
| 17 | Qwen3 VL 235B A22B Instruct | Alibaba Cloud / Qwen Team | 60 |
| 18 | Qwen3 VL 32B Thinking | Alibaba Cloud / Qwen Team | 59 |
| 19 | Qwen3-Next-80B-A3B-Instruct | Alibaba Cloud / Qwen Team | 59 |
| 20 | Qwen3.5-9B | Alibaba Cloud / Qwen Team | 58 |
| 21 | Kimi K2-Instruct-0905 | Moonshot AI | 57 |
| 22 | Kimi K2 Instruct | Moonshot AI | 57 |
| 23 | Qwen3 VL 30B A3B Thinking | Alibaba Cloud / Qwen Team | 56 |
| 24 | Qwen3 VL 32B Instruct | Alibaba Cloud / Qwen Team | 55 |
| 25 | Qwen3 VL 30B A3B Instruct | Alibaba Cloud / Qwen Team | 53 |
| 26 | Qwen3.5-4B | Alibaba Cloud / Qwen Team | 53 |
| 27 | Qwen3 VL 8B Thinking | Alibaba Cloud / Qwen Team | 51 |
| 28 | Qwen3 VL 4B Thinking | Alibaba Cloud / Qwen Team | 47 |
| 29 | Kimi K2 Base | Moonshot AI | 45 |
| 30 | Qwen3 VL 8B Instruct | Alibaba Cloud / Qwen Team | 45 |
| 31 | Qwen3 235B A22B | Alibaba Cloud / Qwen Team | 44 |
| 32 | Qwen3 VL 4B Instruct | Alibaba Cloud / Qwen Team | 40 |
| 33 | Qwen3.5-2B | Alibaba Cloud / Qwen Team | 38 |
| 34 | Qwen3.5-0.8B | Alibaba Cloud / Qwen Team | 21 |
34 of 34 models · score normalized 0–100 where available