Claude Opus 4.8

Adaptive Reasoning, Max Effort

Creator

Anthropic

Released

2026-05-28

Intelligence

55.7

Artificial Analysis Index

Coding

74.3

Artificial Analysis Index

In $/1M

$5.00

input tokens

Out $/1M

$25.00

output tokens

Blended $/1M

$10.00

3:1 blended

Speed

63

tokens / sec

Profile

License, openness, and modality — the governance layer.

License

Proprietary

non-commercial

Weights

Closed

API only

Modalities

Image · Text

Parameters

Claude Opus 4.8 is Anthropic's upgrade to Opus 4.7 and its most capable general-access model at release, with improvements across software engineering, agentic tool use, reasoning, computer use, and knowledge-work benchmarks while shipping at the same price ($5/$25 per million input/output tokens). Performance gains include SWE-Bench Verified (88.6%), SWE-Bench Pro (69.2%), Terminal-Bench 2.1 (74.6%), GPQA Diamond (93.6%), USAMO 2026 (96.7%), Humanity's Last Exam with tools (57.9%), OSWorld-Verified (83.4%), BrowseComp (84.3% single-agent, 88.5% multi-agent), MCP-Atlas (82.2%), and GDPval-AA (1890 Elo). The alignment assessment reports honesty improvements with around a four-fold drop in letting flaws in self-written code pass unremarked, a 17-fold drop relative to Sonnet 4.6 on dishonest agentic code summaries, and broadly improved adherence to Claude's constitution. The model defaults to high effort and exposes new 'extra' (xhigh) and 'max' levels for harder problems. Launches alongside Claude Code dynamic workflows (parallel subagents that plan, execute, and verify codebase-scale migrations), effort control in claude.ai and Cowork, and a Messages API extension that accepts system entries inside the messages array so harnesses can update instructions mid-task without breaking the prompt cache. Fast mode runs at 2.5× speed at $10/$50 per million input/output tokens, three times cheaper than fast mode on previous models. Available across Claude products, the Claude API as `claude-opus-4-8`, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.

Capability profile

Category strength across 20 domains, via LLM Stats.

legal
100finance
100agents
100reasoning
100general
100search
90biology
90physics
90chemistry
90grounding
90
frontend development
90safety
80long context
80spatial reasoning
80code
70math
70vision
70multimodal
70tool calling
70healthcare
60
Via LLM Stats

Benchmark breakdown

Independent evaluation scores, normalized to 0–100.

GPQA Diamond
92
Humanity’s Last Exam
46
MMLU-Pro
SciCode
54
LiveCodeBench
MATH-500
AIME 2025
τ²-Bench (agentic)
94
Terminal-Bench Hard
58
IFBench
62
Via Artificial Analysis

Providers

Inference hosts serving this model — their own pricing and measured performance, cheapest input first.

ProviderIn $/1MOut $/1MThroughputLatencyStatus
Anthropic$5.00$25.00420.50active
All providers

1 host · via LLM Stats

© 2026 NYSGPT2525 LLC