Category
Agentic
Source
Artificial Analysis
evaluation of record
Models covered
—
none ingested
Data status
Demo Purposes Only
invented numbers
Top score
—
not measured
Top model
—
Updated
—
no ingest
A private evaluation developed by Artificial Analysis for frontier agentic capability in long-horizon knowledge work, testing agents on realistic business workflows that require deliverables such as spreadsheets, presentations, and memos.
Leaderboard
The shape of the source chart, drawn on invented numbers. We hold no data for this evaluation.
Demo purposes only · invented numbers · not a measurement
Model names are drawn from our own roster; the scores are invented and measure nothing. See the source for the real leaderboard.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Tests whether an agent can do a junior knowledge worker’s day: long business tasks that have to end in an actual artifact — a spreadsheet, a deck, a memo — not a paragraph describing one. Higher is better, and everything here is hard, because the work is long-horizon and each step can undo the last. The task set is private and held back deliberately, which stops models being trained on it but also means no one outside can audit the tasks or reproduce a score. Read it as a ranking of long-horizon reliability rather than a precise measurement of it.