Artificial Analysis Long Context Reasoning Benchmark Leaderboard

Instruction & Long ContextLive

Category

Instruction & Long Context

Source

Artificial Analysis

evaluation of record

Models covered

479

in our data

Data status

Live

Top score

75.7%

best on record

Top model

GPT-5.2 Codex

OpenAI

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A challenging benchmark measuring language models' ability to extract, reason about, and synthesize information from long-form documents ranging from 10k to 100k tokens (measured using the cl100k_base tokenizer).

Leaderboard

Top 20 of 479 models we hold a score for.

1GPT-5.2 CodexxhighOpenAI
75.7%2GPT-5highOpenAI
75.6%3GPT-5.1highOpenAI
75.0%4Kimi K3Kimi
74.7%5GPT-5.5xhighOpenAI
74.3%6Claude Opus 4.5ReasoningAnthropic
74.0%7GPT-5.3 CodexxhighOpenAI
74.0%8GPT-5.4xhighOpenAI
74.0%9GPT-5.6 LunamaxOpenAI
74.0%10GPT-5.6 TerramaxOpenAI
74.0%11KAT-Coder-Pro V1KwaiKAT
74.0%12MiniMax-M3MiniMax
74.0%13GPT-5.6 SolmaxOpenAI
73.7%14GPT-5.5highOpenAI
73.3%15MiMo-V2.5-ProXiaomi
73.3%16GPT-5mediumOpenAI
72.8%17GPT-5.2xhighOpenAI
72.7%18Gemini 3.1 Pro PreviewGoogle
72.7%19GPT-5.5mediumOpenAI
72.3%20GPT-5.6 TerrahighOpenAI
72.3%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Hands a model documents of 10,000 to 100,000 tokens and asks questions that can only be answered by combining several passages — not by locating one sentence. Higher is better; the score is the share of questions answered correctly. It is deliberately harder than the "needle in a haystack" retrieval tests where models score near-perfectly, which is why a large advertised context window tells you very little about how a model will do here. Performance normally degrades as documents get longer, and a single averaged score hides the point at which a given model starts to break down.

© 2026 NYSGPT2525 LLC