Artificial Analysis Long Context Reasoning Benchmark Leaderboard
Category
Instruction & Long Context
Source
Artificial Analysis
evaluation of record
Models covered
479
in our data
Data status
Live
Top score
75.7%
best on record
Top model
GPT-5.2 Codex
OpenAI
Updated
2026-07-30
last ingest
A challenging benchmark measuring language models' ability to extract, reason about, and synthesize information from long-form documents ranging from 10k to 100k tokens (measured using the cl100k_base tokenizer).
Leaderboard
Top 20 of 479 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Hands a model documents of 10,000 to 100,000 tokens and asks questions that can only be answered by combining several passages — not by locating one sentence. Higher is better; the score is the share of questions answered correctly. It is deliberately harder than the "needle in a haystack" retrieval tests where models score near-perfectly, which is why a large advertised context window tells you very little about how a model will do here. Performance normally degrades as documents get longer, and a single averaged score hides the point at which a given model starts to break down.