Terminal-Bench Hard Benchmark Leaderboard

AgenticLive

Category

Agentic

Source

Artificial Analysis

evaluation of record

Models covered

432

in our data

Data status

Live

Top score

65.9%

best on record

Top model

GPT-5.6 Sol

OpenAI

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

An agentic benchmark evaluating AI capabilities in terminal environments through software engineering, system administration, and data processing tasks.

Leaderboard

Top 20 of 432 models we hold a score for.

1GPT-5.6 SolmaxOpenAI
65.9%2Claude Fable 5Adaptive Reasoning, Max Effort, Opus 4.8 FallbackAnthropic
62.9%3GPT-5.6 SolmediumOpenAI
62.9%4GPT-5.6 TerraxhighOpenAI
62.9%5GPT-5.6 SolhighOpenAI
62.1%6GPT-5.6 SolxhighOpenAI
61.4%7GPT-5.5xhighOpenAI
60.6%8GPT-5.6 SollowOpenAI
60.6%9GPT-5.5highOpenAI
59.8%10Claude Opus 4.8Adaptive Reasoning, Max EffortAnthropic
58.3%11GPT-5.4xhighOpenAI
57.6%12GPT-5.5mediumOpenAI
57.6%13GPT-5.6 TerrahighOpenAI
57.6%14GPT-5.6 TerramaxOpenAI
57.6%15Claude Opus 4.7Non-reasoning, High EffortAnthropic
54.5%16Gemini 3.1 Pro PreviewGoogle
53.8%17Claude Sonnet 4.6Adaptive Reasoning, Max EffortAnthropic
53.0%18GPT-5.3 CodexxhighOpenAI
53.0%19GPT-5.4 minixhighOpenAI
52.3%20GPT-5.5lowOpenAI
52.3%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

The subset of terminal work that current models mostly fail — software engineering, system administration, and data processing carried out through a shell, scored on whether the job actually finished. Higher is better, and the numbers run well below the full Terminal-Bench, which is the reason this cut exists: the parent benchmark is close to saturated. Success is binary per task, so a run that gets 95 percent of the way there scores the same as one that never started, and partial competence is invisible. Expect meaningful variance between runs, since agentic tasks turn on tool calls and timeouts as much as on reasoning.

© 2026 NYSGPT2525 LLC