Creator
LongCat
Released
2026-01-28
Intelligence
17.2
Artificial Analysis Index
Coding
—
Artificial Analysis Index
In $/1M
$0.00
input tokens
Out $/1M
$0.00
output tokens
Blended $/1M
$0.00
3:1 blended
Speed
0
tokens / sec
Profile
License, openness, and modality — the governance layer.
License
MIT
commercial OK
Weights
Open
downloadable
Modalities
Text
Parameters
69B
total
LongCat-Flash-Lite is a lightweight MoE model from Meituan with 68.5B total parameters and only 2.9B-4.5B activated per token. It explores N-gram embedding expansion as a new scaling direction, supporting 256K context length via YaRN. Optimized for agent tooling and programming tasks, achieving 500-700 tokens per second inference speed while maintaining strong performance on coding, math, and agentic benchmarks.
Capability profile
Category strength across 20 domains, via LLM Stats.
Benchmark breakdown
Independent evaluation scores, normalized to 0–100.
Providers
Inference hosts serving this model — their own pricing and measured performance, cheapest input first.
| Provider | In $/1M | Out $/1M | Throughput | Latency | Status |
|---|---|---|---|---|---|
| Meituan | $0.10 | $0.40 | 500 | 1.50 | active |
1 host · via LLM Stats