Humanity's Last Exam
A benchmark of expert-level academic questions to assess AI capabilities
Models scored
91
evaluated
Modality
multimodal
Category
math
+2 more
Published
2025
arxiv.org
Citations
337
Semantic Scholar
Influential
40
citations
References
61
cited works
Venue
Nature
published in
Abstract
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, et al. (+496)
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve more than 90% accuracy on popular benchmarks such as Measuring Massive Multitask Language Understanding1, limiting informed measurement of state-of-the-art LLM capabilities. Here, in response, we introduce Humanity’s Last Exam (HLE), a multi-modal benchmark at the frontier of human knowledge, designed to be an expert-level closed-ended academic benchmark with broad subject coverage. HLE consists of 2,500 questions across dozens of subjects, including mathematics, humanities and the natural sciences. HLE is developed globally by subject-matter experts and consists of multiple-choice and short-answer questions suitable for automated grading. Each question has a known solution that is unambiguous and easily verifiable but cannot be quickly answered by internet retrieval. State-of-the-art LLMs demonstrate low accuracy and calibration on HLE, highlighting a marked gap between current LLM capabilities and the expert human frontier on closed-ended academic questions. To inform research and policymaking upon a clear understanding of model capabilities, we publicly release HLE at https://lastexam.ai. Humanity’s Last Exam, a multi-modal benchmark at the frontier of human knowledge, is designed to be an expert-level closed-ended academic benchmark with broad subject coverage.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Claude Mythos Preview | Anthropic | 65 |
| 02 | Claude Opus 5 | Anthropic | 65 |
| 03 | Claude Fable 5 | Anthropic | 65 |
| 04 | Muse Spark 1.1 | Meta | 62 |
| 05 | Muse Spark | Meta | 58 |
| 06 | Claude Opus 4.8 | Anthropic | 58 |
| 07 | Claude Sonnet 5 | Anthropic | 57 |
| 08 | GPT-5.5 Pro | OpenAI | 57 |
| 09 | Kimi K3 | Moonshot AI | 56 |
| 10 | Seed 2.1 Pro | ByteDance | 56 |
| 11 | Claude Opus 4.7 | Anthropic | 55 |
| 12 | GLM-5.2 | Zhipu AI | 55 |
| 13 | Seed 2.1 Turbo | ByteDance | 55 |
| 14 | Claude Opus 4.6 | Anthropic | 53 |
| 15 | GLM-5.1 | Zhipu AI | 52 |
| 16 | GPT-5.5 | OpenAI | 52 |
| 17 | Gemini 3.1 Pro | 51 | |
| 18 | Kimi K2-Thinking-0905 | Moonshot AI | 51 |
| 19 | Grok-4 Heavy | xAI | 51 |
| 20 | Kimi K2.5 | Moonshot AI | 50 |
| 21 | Claude Sonnet 4.6 | Anthropic | 49 |
| 22 | Qwen3.5-27B | Alibaba Cloud / Qwen Team | 49 |
| 23 | DeepSeek-V4-Pro-Max | DeepSeek | 48 |
| 24 | Qwen3.5-122B-A10B | Alibaba Cloud / Qwen Team | 48 |
| 25 | Qwen3.5-35B-A3B | Alibaba Cloud / Qwen Team | 47 |
| 26 | Gemini 3 Pro | 46 | |
| 27 | DeepSeek-V4-Flash-Max | DeepSeek | 45 |
| 28 | Gemini 3 Flash | 44 | |
| 29 | GLM-4.7 | Zhipu AI | 43 |
| 30 | Qwen3.7 Max | Alibaba Cloud / Qwen Team | 41 |
| 31 | DeepSeek-V3.2 | DeepSeek | 41 |
| 32 | Gemini 3.5 Flash | 40 | |
| 33 | Grok-4 | xAI | 40 |
| 34 | GPT-5.4 | OpenAI | 40 |
| 35 | ERNIE 5.0 | Baidu | 39 |
| 36 | Nemotron 3 Ultra (550B A55B) | NVIDIA | 37 |
| 37 | GPT-5.2 Pro | OpenAI | 37 |
| 38 | Kimi K2.6 | Moonshot AI | 36 |
| 39 | Qwen3.7-Plus | Alibaba Cloud / Qwen Team | 35 |
| 40 | GPT-5.2 | OpenAI | 35 |
| 41 | MiMo-V2.5-Pro | Xiaomi | 34 |
| 42 | DeepSeek-V3.2-Speciale | DeepSeek | 31 |
| 43 | Qwen3.6 Plus | Alibaba Cloud / Qwen Team | 29 |
| 44 | Qwen3.5-397B-A17B | Alibaba Cloud / Qwen Team | 29 |
| 45 | GPT-5.4 mini | OpenAI | 28 |
| 46 | Gemma 4 31B | 27 | |
| 47 | LongCat-Flash-Thinking-2601 | Meituan | 25 |
| 48 | DeepSeek-V3.2 (Thinking) | DeepSeek | 25 |
| 49 | GPT-5 | OpenAI | 25 |
| 50 | GPT-5.4 nano | OpenAI | 24 |
| 51 | Qwen3.6-27B | Alibaba Cloud / Qwen Team | 24 |
| 52 | Nemotron 3 Super (120B A12B) | NVIDIA | 23 |
| 53 | MiMo-V2-Flash | Xiaomi | 22 |
| 54 | MiniMax M2.1 | MiniMax | 22 |
| 55 | Gemini 2.5 Pro Preview 06-05 | 22 | |
| 56 | Qwen3.6-35B-A3B | Alibaba Cloud / Qwen Team | 21 |
| 57 | Grok 4 Fast | xAI | 20 |
| 58 | DeepSeek-V3.2-Exp | DeepSeek | 20 |
| 59 | Qwen3-235B-A22B-Thinking-2507 | Alibaba Cloud / Qwen Team | 18 |
| 60 | MAI-Code-1-Flash | Microsoft | 18 |
60 of 91 models · score normalized 0–100 where available