Gemini 3.5 Flash-Lite
Creator
Google DeepMind
Released
2026-07-21
Intelligence
36.5
Artificial Analysis Index
Coding
49.3
Artificial Analysis Index
In $/1M
$0.30
input tokens
Out $/1M
$2.50
output tokens
Blended $/1M
$0.85
3:1 blended
Speed
345
tokens / sec
Profile
License, openness, and modality — the governance layer.
License
Proprietary
non-commercial
Weights
Closed
API only
Modalities
Audio · Image · Text · Video
Parameters
—
Gemini 3.5 Flash-Lite is Google's low-latency, cost-effective multimodal reasoning model for high-throughput agentic workflows, document processing, data extraction, translation, and classification. It supports text, image, video, audio, and PDF inputs, a 1 million-token context window, and a 65,536-token text output.
Capability profile
Category strength across 20 domains, via LLM Stats.
Benchmark breakdown
Independent evaluation scores, normalized to 0–100.
Providers
Inference hosts serving this model — their own pricing and measured performance, cheapest input first.
| Provider | In $/1M | Out $/1M | Throughput | Latency | Status |
|---|---|---|---|---|---|
| $0.30 | $2.50 | — | — | active |
1 host · via LLM Stats