Creator

Inception

Released

2026-02-20

Intelligence

21.4

Artificial Analysis Index

Coding

31.1

Artificial Analysis Index

In $/1M

$0.25

input tokens

Out $/1M

$0.75

output tokens

Blended $/1M

$0.38

3:1 blended

Speed

905

tokens / sec

Profile

License, openness, and modality — the governance layer.

License

Proprietary

non-commercial

Weights

Closed

API only

Modalities

Text

Parameters

Mercury 2 is the fastest reasoning LLM, built on diffusion-based language model (dLLM) architecture. Instead of generating text token-by-token, it refines multiple text blocks simultaneously, achieving over 1,000 tokens per second on Nvidia Blackwell GPUs — 5x faster than leading speed-optimized LLMs. Supports tool usage and JSON output with 128K context window.

Capability profile

Category strength across 20 domains, via LLM Stats.

general
70instruction following
70math
60biology
60physics
60
chemistry
60reasoning
60code
50tool calling
50communication
50
Via LLM Stats

Benchmark breakdown

Independent evaluation scores, normalized to 0–100.

GPQA Diamond
77
Humanity’s Last Exam
16
MMLU-Pro
SciCode
39
LiveCodeBench
MATH-500
AIME 2025
τ²-Bench (agentic)
71
Terminal-Bench Hard
27
IFBench
70
Via Artificial Analysis

Providers

Inference hosts serving this model — their own pricing and measured performance, cheapest input first.

ProviderIn $/1MOut $/1MThroughputLatencyStatus
Inception$0.25$0.751,0091.70active
All providers

1 host · via LLM Stats

© 2026 NYSGPT2525 LLC