Harvey LAB-AA Benchmark Leaderboard

AgenticDemo Purposes Only

Category

Agentic

Source

Artificial Analysis

evaluation of record

Models covered

none ingested

Data status

Demo Purposes Only

invented numbers

Top score

not measured

Top model

Updated

no ingest

Full results

artificialanalysis.ai

on the source

Artificial Analysis' implementation of Harvey's Legal Agent Benchmark (LAB), testing AI agents on real-world legal work from Harvey's dataset of 120 private tasks spanning 24 legal practice areas. The agent reads case documents in a sandbox and produces legal deliverables (e.g., memos, disclosure schedules, deposition summaries), graded criterion-by-criterion by a single LLM rubric judge.

Leaderboard

The shape of the source chart, drawn on invented numbers. We hold no data for this evaluation.

Demo purposes only · invented numbers · not a measurement

1Claude Opus 5Adaptive Reasoning, Max EffortAnthropic
55.0%2GPT-5.6 SolmaxOpenAI
50.2%3Kimi K3Kimi
45.2%4Grok 4.5highSpaceXAI
39.5%5GLM-5.2maxZ AI
36.2%6Muse Spark 1.1xhighMeta
32.9%7Gemini 3.5 FlashhighGoogle
28.7%8Qwen3.7 MaxAlibaba
26.5%9MiniMax-M3MiniMax
24.2%10DeepSeek V4 ProReasoning, Max EffortDeepSeek
21.8%11Motif 3BetaMotif Technologies
19.4%12MiMo-V2.5-ProXiaomi
17.5%13Hy3Tencent
16.2%14Nex-N2-ProNex AGI
14.5%15InklingxhighThinking Machines
13.3%16Agnes 2.5 Pro AlphaSapiens AI
11.7%17JT-4.1 Flash 236B A21BChina Mobile
10.3%18Nemotron 3 Ultra 550B A55BReasoningNVIDIA
9.1%19Claude Opus 5Adaptive Reasoning, Xhigh EffortAnthropic
8.3%20Claude Fable 5Adaptive Reasoning, Max Effort, Opus 4.8 FallbackAnthropic
7.3%

Model names are drawn from our own roster; the scores are invented and measure nothing. See the source for the real leaderboard.

Source

Plain explanation

What it measures, how to read the number, and what to watch out for.

Puts models on genuine legal work — 120 private tasks across 24 practice areas — where the agent reads case documents in a sandbox and produces a deliverable a lawyer would recognize: a memo, a disclosure schedule, a deposition summary. Grading runs criterion by criterion against a rubric, so the score is the share of required elements the deliverable actually contained; higher is better. One language model does all the grading, which is consistent and cheap but inherits that model’s blind spots, and a rubric rewards completeness more readily than judgment. The task set is private, which prevents training on it and also means the scores cannot be independently reproduced.

© 2026 NYSGPT2525 LLC