Creator
DeepSeek
Released
2025-03-25
Intelligence
15.4
Artificial Analysis Index
Coding
21.2
Artificial Analysis Index
In $/1M
$0.27
input tokens
Out $/1M
$1.12
output tokens
Blended $/1M
$0.48
3:1 blended
Speed
0
tokens / sec
Profile
License, openness, and modality — the governance layer.
License
MIT + Model License (Commercial use allowed)
non-commercial
Weights
Open
downloadable
Modalities
Text
Parameters
671B
total
A powerful Mixture-of-Experts (MoE) language model with 671B total parameters (37B activated per token). Features Multi-head Latent Attention (MLA), auxiliary-loss-free load balancing, and multi-token prediction training. Pre-trained on 14.8T tokens with strong performance in reasoning, math, and code tasks.
Capability profile
Category strength across 20 domains, via LLM Stats.
Benchmark breakdown
Independent evaluation scores, normalized to 0–100.