Creator
Alibaba
Released
2025-09-11
Intelligence
13.7
Artificial Analysis Index
Coding
—
Artificial Analysis Index
In $/1M
$0.50
input tokens
Out $/1M
$2.00
output tokens
Blended $/1M
$0.88
3:1 blended
Speed
196
tokens / sec
Profile
License, openness, and modality — the governance layer.
License
Apache 2.0
commercial OK
Weights
Open
downloadable
Modalities
Text
Parameters
80B
total
Qwen3-Next-80B-A3B-Instruct is the first in the Qwen3-Next series, featuring groundbreaking architectural innovations. It uses Hybrid Attention combining Gated DeltaNet and Gated Attention for efficient ultra-long context modeling, High-Sparsity MoE with 512 experts (10 activated + 1 shared) achieving extreme low activation ratio, and Multi-Token Prediction for improved performance and faster inference. With 80B total parameters and only 3B activated, it outperforms Qwen3-32B-Base with 10% training cost and 10x throughput for 32K+ contexts. The model performs on par with Qwen3-235B-A22B-Instruct-2507 while excelling at ultra-long-context tasks up to 256K tokens (extensible to 1M with YaRN). Architecture: 48 layers, 15T training tokens, hybrid layout of 12*(3*(Gated DeltaNet->MoE)->(Gated Attention->MoE)).
Capability profile
Category strength across 20 domains, via LLM Stats.
Benchmark breakdown
Independent evaluation scores, normalized to 0–100.