Category
Agentic
Source
Artificial Analysis
evaluation of record
Models covered
178
in our data
Data status
Live
Top score
89.5%
best on record
Top model
GPT-5.6 Sol
OpenAI
Updated
2026-07-30
last ingest
A verified refresh of Terminal-Bench v2.0 — 89 curated tasks across software engineering, system administration, data processing, model training, and security, with environment and instruction fixes so scores reflect agent capability rather than environment gaps.
Leaderboard
Top 20 of 178 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Puts a model in a real terminal with 89 hand-checked jobs — build something, administer a system, process data, train a model, close a security hole — and scores whether the job actually completed. Higher is better; the frontier is now near 90 percent. Version 2.1 exists because earlier versions punished models for broken environments rather than bad reasoning, and the fixes lifted scores across the board, so v2.1 numbers are not comparable with v2.0 or v1. As the leaders approach 90 percent the remaining gap between them is a handful of tasks, well inside run-to-run noise for agentic work.