Humanity's Last Exam

A benchmark of expert-level academic questions to assess AI capabilities

Models scored

91

evaluated

Modality

multimodal

Category

math

+2 more

Published

2025

arxiv.org

Citations

337

Semantic Scholar

Influential

40

citations

References

61

cited works

Venue

Nature

published in

Abstract

Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, et al. (+496)

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve more than 90% accuracy on popular benchmarks such as Measuring Massive Multitask Language Understanding1, limiting informed measurement of state-of-the-art LLM capabilities. Here, in response, we introduce Humanity’s Last Exam (HLE), a multi-modal benchmark at the frontier of human knowledge, designed to be an expert-level closed-ended academic benchmark with broad subject coverage. HLE consists of 2,500 questions across dozens of subjects, including mathematics, humanities and the natural sciences. HLE is developed globally by subject-matter experts and consists of multiple-choice and short-answer questions suitable for automated grading. Each question has a known solution that is unambiguous and easily verifiable but cannot be quickly answered by internet retrieval. State-of-the-art LLMs demonstrate low accuracy and calibration on HLE, highlighting a marked gap between current LLM capabilities and the expert human frontier on closed-ended academic questions. To inform research and policymaking upon a clear understanding of model capabilities, we publicly release HLE at https://lastexam.ai. Humanity’s Last Exam, a multi-modal benchmark at the frontier of human knowledge, is designed to be an expert-level closed-ended academic benchmark with broad subject coverage.

mathreasoningvision

Search

#ModelLabScore
01Claude Mythos PreviewAnthropic65
02Claude Opus 5Anthropic65
03Claude Fable 5Anthropic65
04Muse Spark 1.1Meta62
05Muse SparkMeta58
06Claude Opus 4.8Anthropic58
07Claude Sonnet 5Anthropic57
08GPT-5.5 ProOpenAI57
09Kimi K3Moonshot AI56
10Seed 2.1 ProByteDance56
11Claude Opus 4.7Anthropic55
12GLM-5.2Zhipu AI55
13Seed 2.1 TurboByteDance55
14Claude Opus 4.6Anthropic53
15GLM-5.1Zhipu AI52
16GPT-5.5OpenAI52
17Gemini 3.1 ProGoogle51
18Kimi K2-Thinking-0905Moonshot AI51
19Grok-4 HeavyxAI51
20Kimi K2.5Moonshot AI50
21Claude Sonnet 4.6Anthropic49
22Qwen3.5-27BAlibaba Cloud / Qwen Team49
23DeepSeek-V4-Pro-MaxDeepSeek48
24Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team48
25Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team47
26Gemini 3 ProGoogle46
27DeepSeek-V4-Flash-MaxDeepSeek45
28Gemini 3 FlashGoogle44
29GLM-4.7Zhipu AI43
30Qwen3.7 MaxAlibaba Cloud / Qwen Team41
31DeepSeek-V3.2DeepSeek41
32Gemini 3.5 FlashGoogle40
33Grok-4xAI40
34GPT-5.4OpenAI40
35ERNIE 5.0Baidu39
36Nemotron 3 Ultra (550B A55B)NVIDIA37
37GPT-5.2 ProOpenAI37
38Kimi K2.6Moonshot AI36
39Qwen3.7-PlusAlibaba Cloud / Qwen Team35
40GPT-5.2OpenAI35
41MiMo-V2.5-ProXiaomi34
42DeepSeek-V3.2-SpecialeDeepSeek31
43Qwen3.6 PlusAlibaba Cloud / Qwen Team29
44Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team29
45GPT-5.4 miniOpenAI28
46Gemma 4 31BGoogle27
47LongCat-Flash-Thinking-2601Meituan25
48DeepSeek-V3.2 (Thinking)DeepSeek25
49GPT-5OpenAI25
50GPT-5.4 nanoOpenAI24
51Qwen3.6-27BAlibaba Cloud / Qwen Team24
52Nemotron 3 Super (120B A12B)NVIDIA23
53MiMo-V2-FlashXiaomi22
54MiniMax M2.1MiniMax22
55Gemini 2.5 Pro Preview 06-05Google22
56Qwen3.6-35B-A3BAlibaba Cloud / Qwen Team21
57Grok 4 FastxAI20
58DeepSeek-V3.2-ExpDeepSeek20
59Qwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team18
60MAI-Code-1-FlashMicrosoft18

60 of 91 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC