Terminal-Bench v2.1 Benchmark Leaderboard

AgenticLive

Category

Agentic

Source

Artificial Analysis

evaluation of record

Models covered

178

in our data

Data status

Live

Top score

89.5%

best on record

Top model

GPT-5.6 Sol

OpenAI

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A verified refresh of Terminal-Bench v2.0 — 89 curated tasks across software engineering, system administration, data processing, model training, and security, with environment and instruction fixes so scores reflect agent capability rather than environment gaps.

Leaderboard

Top 20 of 178 models we hold a score for.

1GPT-5.6 SolxhighOpenAI
89.5%2Claude Opus 5Adaptive Reasoning, Max EffortAnthropic
89.1%3Claude Opus 5Adaptive Reasoning, Xhigh EffortAnthropic
88.0%4GPT-5.6 SolmaxOpenAI
88.0%5GPT-5.6 TerramaxOpenAI
88.0%6Claude Opus 5Adaptive Reasoning, High EffortAnthropic
87.6%7GPT-5.6 SolhighOpenAI
87.3%8Claude Opus 5Adaptive Reasoning, Medium EffortAnthropic
86.1%9GPT-5.6 SolmediumOpenAI
86.1%10Kimi K3Kimi
85.0%11Claude Fable 5Adaptive Reasoning, Max Effort, Opus 4.8 FallbackAnthropic
84.6%12Claude Opus 4.8Adaptive Reasoning, Max EffortAnthropic
84.6%13GPT-5.5xhighOpenAI
84.3%14Claude Opus 4.7Adaptive Reasoning, Max EffortAnthropic
83.1%15Grok 4.5highSpaceXAI
81.6%16GPT-5.6 LunamaxOpenAI
80.9%17Claude Sonnet 5Adaptive Reasoning, Max EffortAnthropic
80.5%18GPT-5.5mediumOpenAI
80.5%19GPT-5.6 TerraxhighOpenAI
80.1%20GPT-5.5highOpenAI
79.4%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Puts a model in a real terminal with 89 hand-checked jobs — build something, administer a system, process data, train a model, close a security hole — and scores whether the job actually completed. Higher is better; the frontier is now near 90 percent. Version 2.1 exists because earlier versions punished models for broken environments rather than bad reasoning, and the fixes lifted scores across the board, so v2.1 numbers are not comparable with v2.0 or v1. As the leaders approach 90 percent the remaining gap between them is a handful of tasks, well inside run-to-run noise for agentic work.

© 2026 NYSGPT2525 LLC