Qwen3 Next 80B A3B Instruct

Creator

Alibaba

Released

2025-09-11

Intelligence

13.7

Artificial Analysis Index

Coding

Artificial Analysis Index

In $/1M

$0.50

input tokens

Out $/1M

$2.00

output tokens

Blended $/1M

$0.88

3:1 blended

Speed

196

tokens / sec

Profile

License, openness, and modality — the governance layer.

License

Apache 2.0

commercial OK

Weights

Open

downloadable

Modalities

Text

Parameters

80B

total

Qwen3-Next-80B-A3B-Instruct is the first in the Qwen3-Next series, featuring groundbreaking architectural innovations. It uses Hybrid Attention combining Gated DeltaNet and Gated Attention for efficient ultra-long context modeling, High-Sparsity MoE with 512 experts (10 activated + 1 shared) achieving extreme low activation ratio, and Multi-Token Prediction for improved performance and faster inference. With 80B total parameters and only 3B activated, it outperforms Qwen3-32B-Base with 10% training cost and 10x throughput for 32K+ contexts. The model performs on par with Qwen3-235B-A22B-Instruct-2507 while excelling at ultra-long-context tasks up to 256K tokens (extensible to 1M with YaRN). Architecture: 48 layers, 15T training tokens, hybrid layout of 12*(3*(Gated DeltaNet->MoE)->(Gated Attention->MoE)).

Capability profile

Category strength across 20 domains, via LLM Stats.

writing
90creativity
90legal
80language
80structured output
80instruction following
80math
70agents
70biology
70finance
70general
70
physics
70chemistry
70healthcare
70economics
60reasoning
60code
50vision
50multimodal
50tool calling
50communication
50spatial reasoning
50
Via LLM Stats

Benchmark breakdown

Independent evaluation scores, normalized to 0–100.

GPQA Diamond
74
Humanity’s Last Exam
7
MMLU-Pro
82
SciCode
31
LiveCodeBench
68
MATH-500
AIME 2025
66
τ²-Bench (agentic)
22
Terminal-Bench Hard
8
IFBench
40
Via Artificial Analysis
© 2026 NYSGPT2525 LLC