Category
Agentic
Source
Artificial Analysis
evaluation of record
Models covered
432
in our data
Data status
Live
Top score
65.9%
best on record
Top model
GPT-5.6 Sol
OpenAI
Updated
2026-07-30
last ingest
An agentic benchmark evaluating AI capabilities in terminal environments through software engineering, system administration, and data processing tasks.
Leaderboard
Top 20 of 432 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
The subset of terminal work that current models mostly fail — software engineering, system administration, and data processing carried out through a shell, scored on whether the job actually finished. Higher is better, and the numbers run well below the full Terminal-Bench, which is the reason this cut exists: the parent benchmark is close to saturated. Success is binary per task, so a run that gets 95 percent of the way there scores the same as one that never started, and partial competence is invisible. Expect meaningful variance between runs, since agentic tasks turn on tool calls and timeouts as much as on reasoning.