Gemini 3.1 Flash-Lite
Creator
Google DeepMind
Released
2026-03-03
Intelligence
25.0
Artificial Analysis Index
Coding
34.7
Artificial Analysis Index
In $/1M
$0.25
input tokens
Out $/1M
$1.50
output tokens
Blended $/1M
$0.56
3:1 blended
Speed
291
tokens / sec
Profile
License, openness, and modality — the governance layer.
License
Proprietary
non-commercial
Weights
Closed
API only
Modalities
Audio · Image · Text · Video
Parameters
—
Gemini 3.1 Flash-Lite is the first Flash-Lite model in the Gemini 3 series. It is optimized for high-volume, latency-sensitive tasks like translation, content moderation, and classification. It delivers enhanced performance at a fraction of the cost of larger models, with 2.5x faster Time to First Answer Token and 45% increased output speed compared to 2.5 Flash. Supports text, image, video, audio, and PDF input with a 1 million-token context window.
Capability profile
Category strength across 20 domains, via LLM Stats.
Benchmark breakdown
Independent evaluation scores, normalized to 0–100.
Providers
Inference hosts serving this model — their own pricing and measured performance, cheapest input first.
| Provider | In $/1M | Out $/1M | Throughput | Latency | Status |
|---|---|---|---|---|---|
| $0.25 | $1.50 | — | — | active |
1 host · via LLM Stats