Category
Agentic
Source
Artificial Analysis
evaluation of record
Models covered
7
in our data
Data status
Live
Top score
30.8%
best on record
Top model
Kimi K3
Moonshot AI
Updated
2026-07-30
last ingest
A benchmark measuring agentic task completion across simulated SaaS application environments, scoring the share of each task's objectives completed without guardrail violations.
Leaderboard
Top 7 of 7 models we hold a score for.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Scores agents on automating work inside simulated SaaS tools, counting how much of each task got finished and discounting any run that broke a guardrail on the way. Higher is better; scores are a share of objectives, and leaders currently sit around 30 percent. The guardrail condition is what makes this different from a plain capability test: a model that completes the task by taking an unsafe shortcut scores worse than one that stops, so it measures restraint as much as competence. The environments are simulated, so a score is evidence about behavior in a sandbox rather than a guarantee about a live production account.