Category
Agentic
Source
Artificial Analysis
evaluation of record
Models covered
—
none ingested
Data status
Demo Purposes Only
invented numbers
Top score
—
not measured
Top model
—
Updated
—
no ingest
Artificial Analysis' implementation of Harvey's Legal Agent Benchmark (LAB), testing AI agents on real-world legal work from Harvey's dataset of 120 private tasks spanning 24 legal practice areas. The agent reads case documents in a sandbox and produces legal deliverables (e.g., memos, disclosure schedules, deposition summaries), graded criterion-by-criterion by a single LLM rubric judge.
Leaderboard
The shape of the source chart, drawn on invented numbers. We hold no data for this evaluation.
Demo purposes only · invented numbers · not a measurement
Model names are drawn from our own roster; the scores are invented and measure nothing. See the source for the real leaderboard.
Plain explanation
What it measures, how to read the number, and what to watch out for.
Puts models on genuine legal work — 120 private tasks across 24 practice areas — where the agent reads case documents in a sandbox and produces a deliverable a lawyer would recognize: a memo, a disclosure schedule, a deposition summary. Grading runs criterion by criterion against a rubric, so the score is the share of required elements the deliverable actually contained; higher is better. One language model does all the grading, which is consistent and cheap but inherits that model’s blind spots, and a rubric rewards completeness more readily than judgment. The task set is private, which prevents training on it and also means the scores cannot be independently reproduced.