AA-Briefcase: Agentic Knowledge Work Benchmark

AgenticDemo Purposes Only

Category

Agentic

Source

Artificial Analysis

evaluation of record

Models covered

none ingested

Data status

Demo Purposes Only

invented numbers

Top score

not measured

Top model

Updated

no ingest

Full results

artificialanalysis.ai

on the source

A private evaluation developed by Artificial Analysis for frontier agentic capability in long-horizon knowledge work, testing agents on realistic business workflows that require deliverables such as spreadsheets, presentations, and memos.

Leaderboard

The shape of the source chart, drawn on invented numbers. We hold no data for this evaluation.

Demo purposes only · invented numbers · not a measurement

1Claude Opus 5Adaptive Reasoning, Max EffortAnthropic
1,5482GPT-5.6 SolmaxOpenAI
1,5013Kimi K3Kimi
1,4804Grok 4.5highSpaceXAI
1,4335GLM-5.2maxZ AI
1,4076Muse Spark 1.1xhighMeta
1,3507Gemini 3.5 FlashhighGoogle
1,3288Qwen3.7 MaxAlibaba
1,2939MiniMax-M3MiniMax
1,23610DeepSeek V4 ProReasoning, Max EffortDeepSeek
1,19611Motif 3BetaMotif Technologies
1,15712MiMo-V2.5-ProXiaomi
1,10413Hy3Tencent
1,07414Nex-N2-ProNex AGI
1,03615InklingxhighThinking Machines
99116Agnes 2.5 Pro AlphaSapiens AI
95817JT-4.1 Flash 236B A21BChina Mobile
93718Nemotron 3 Ultra 550B A55BReasoningNVIDIA
88619Claude Opus 5Adaptive Reasoning, Xhigh EffortAnthropic
86420Claude Fable 5Adaptive Reasoning, Max Effort, Opus 4.8 FallbackAnthropic
803

Model names are drawn from our own roster; the scores are invented and measure nothing. See the source for the real leaderboard.

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Tests whether an agent can do a junior knowledge worker’s day: long business tasks that have to end in an actual artifact — a spreadsheet, a deck, a memo — not a paragraph describing one. Higher is better, and everything here is hard, because the work is long-horizon and each step can undo the last. The task set is private and held back deliberately, which stops models being trained on it but also means no one outside can audit the tasks or reproduce a score. Read it as a ranking of long-horizon reliability rather than a precise measurement of it.

© 2026 NYSGPT2525 LLC