Claude Sonnet 5

Adaptive Reasoning, Max Effort

Creator

Anthropic

Released

2026-06-30

Intelligence

53.4

Artificial Analysis Index

Coding

71.5

Artificial Analysis Index

In $/1M

$2.00

input tokens

Out $/1M

$10.00

output tokens

Blended $/1M

$4.00

3:1 blended

Speed

91

tokens / sec

Profile

License, openness, and modality — the governance layer.

License

Proprietary

non-commercial

Weights

Closed

API only

Modalities

Image · Text

Parameters

Claude Sonnet 5 is Anthropic's most agentic Sonnet-class model, an upgrade to Sonnet 4.6 that narrows the gap to Opus 4.8 on reasoning, tool use, coding, computer use, and knowledge work while staying lower priced. It plans, uses tools like browsers and terminals, and runs autonomously for long-horizon tasks. Capability gains include SWE-Bench Verified (85.2%), SWE-Bench Pro (63.2%), SWE-Bench Multilingual (78.3%), Terminal-Bench 2.1 (80.4%), OSWorld-Verified (81.2%), BrowseComp (84.7% single-agent, 86.6% multi-agent), Humanity's Last Exam with tools (57.4%), USAMO 2026 (79.5%), GDPval-AA v2 (1618 Elo), HealthBench Professional (57.8%), and FrontierCode v1 (38.8%). It supports adaptive thinking with selectable effort levels up to 'extra high' (xhigh) and a 1M-token context window with context compaction. The safety assessment found lower rates of misaligned behavior, hallucination, and sycophancy than Sonnet 4.6, with improved prompt-injection robustness; it ships with cyber safeguards enabled by default and uses an updated tokenizer (input maps to roughly 1.0-1.35x more tokens than Sonnet 4.6). Default model on Free and Pro plans and available to Max, Team, and Enterprise users, in Claude Code, and on the Claude Platform. Launches with introductory pricing of $2/$10 per million input/output tokens through August 31, 2026, then $3/$15. Available via the Claude API as `claude-sonnet-5`.

Capability profile

Category strength across 20 domains, via LLM Stats.

finance
100legal
100agents
100reasoning
100general
100math
100frontend development
90
search
80code
60vision
60healthcare
60multimodal
60tool calling
50
Via LLM Stats

Benchmark breakdown

Independent evaluation scores, normalized to 0–100.

GPQA Diamond
91
Humanity’s Last Exam
40
MMLU-Pro
SciCode
54
LiveCodeBench
MATH-500
AIME 2025
τ²-Bench (agentic)
Terminal-Bench Hard
IFBench
Via Artificial Analysis

Providers

Inference hosts serving this model — their own pricing and measured performance, cheapest input first.

ProviderIn $/1MOut $/1MThroughputLatencyStatus
Anthropic$3.00$15.00420.40active
All providers

1 host · via LLM Stats

© 2026 NYSGPT2525 LLC