Humanity's Last Exam Benchmark Leaderboard

Knowledge & ReasoningLive

Category

Knowledge & Reasoning

Source

Artificial Analysis

evaluation of record

Models covered

547

in our data

Data status

Live

Top score

53.3%

best on record

Top model

Claude Fable 5

Anthropic

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A frontier-level benchmark with 2,500 expert-vetted questions across mathematics, sciences, and humanities, designed to be the final closed-ended academic evaluation.

Leaderboard

Top 20 of 547 models we hold a score for.

1Claude Fable 5Adaptive Reasoning, Max Effort, Opus 4.8 FallbackAnthropic
53.3%2Claude Opus 5Adaptive Reasoning, Max EffortAnthropic
52.6%3Claude Opus 5Adaptive Reasoning, Xhigh EffortAnthropic
52.5%4Claude Opus 5Adaptive Reasoning, High EffortAnthropic
50.7%5Claude Opus 5Adaptive Reasoning, Medium EffortAnthropic
49.2%6GPT-5.6 SolmaxOpenAI
47.2%7Claude Opus 4.8Adaptive Reasoning, Max EffortAnthropic
45.7%8Muse Spark 1.1xhighMeta
45.1%9GPT-5.6 SolxhighOpenAI
44.7%10Gemini 3.1 Pro PreviewGoogle
44.7%11GPT-5.5xhighOpenAI
44.3%12Kimi K3Kimi
44.3%13GPT-5.6 SolhighOpenAI
44.1%14GPT-5.5highOpenAI
43.0%15GPT-5.6 TerramaxOpenAI
41.8%16GPT-5.4xhighOpenAI
41.6%17Claude Opus 5Adaptive Reasoning, Low EffortAnthropic
41.3%18Gemini 3.5 FlashhighGoogle
41.0%19GPT-5.5mediumOpenAI
40.6%20Grok 4.5highSpaceXAI
40.3%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Two and a half thousand questions written and vetted by specialists to sit at the edge of what an expert in the field could answer, spanning mathematics, the sciences, and the humanities. Higher is better; the frontier is now around half, up from single digits two years ago, so 10-point gaps here still separate genuine capability tiers rather than noise. Scores depend heavily on whether a model was allowed tools and search — a tools-enabled run and a closed-book run are different measurements that get quoted interchangeably, so always check which one a number is. It was built to be the last closed-ended academic exam worth running, which means it will saturate, probably sooner than its authors intended.

© 2026 NYSGPT2525 LLC