AutomationBench-AA: Agentic SaaS Workflow Benchmark

AgenticLive

Category

Agentic

Source

Artificial Analysis

evaluation of record

Models covered

7

in our data

Data status

Live

Top score

30.8%

best on record

Top model

Kimi K3

Moonshot AI

Updated

2026-07-30

last ingest

Full results

artificialanalysis.ai

on the source

A benchmark measuring agentic task completion across simulated SaaS application environments, scoring the share of each task's objectives completed without guardrail violations.

Leaderboard

Top 7 of 7 models we hold a score for.

1Kimi K3Moonshot AI
30.8%2Claude Opus 5Anthropic
26.0%3GPT-5.6 SolOpenAI
18.1%4Claude Fable 5Anthropic
17.4%5GPT-5.6 TerraOpenAI
15.2%6GPT-5.6 LunaOpenAI
14.9%7Claude Sonnet 5Anthropic
13.5%

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Scores agents on automating work inside simulated SaaS tools, counting how much of each task got finished and discounting any run that broke a guardrail on the way. Higher is better; scores are a share of objectives, and leaders currently sit around 30 percent. The guardrail condition is what makes this different from a plain capability test: a model that completes the task by taking an unsafe shortcut scores worse than one that stops, so it measures restraint as much as competence. The environments are simulated, so a score is evidence about behavior in a sandbox rather than a guarantee about a live production account.

© 2026 NYSGPT2525 LLC