Gemini 3.1 Flash-Lite

Creator

Google DeepMind

Released

2026-03-03

Intelligence

25.0

Artificial Analysis Index

Coding

34.7

Artificial Analysis Index

In $/1M

$0.25

input tokens

Out $/1M

$1.50

output tokens

Blended $/1M

$0.56

3:1 blended

Speed

291

tokens / sec

Profile

License, openness, and modality — the governance layer.

License

Proprietary

non-commercial

Weights

Closed

API only

Modalities

Audio · Image · Text · Video

Parameters

Gemini 3.1 Flash-Lite is the first Flash-Lite model in the Gemini 3 series. It is optimized for high-volume, latency-sensitive tasks like translation, content moderation, and classification. It delivers enhanced performance at a fraction of the cost of larger models, with 2.5x faster Time to First Answer Token and 45% increased output speed compared to 2.5 Flash. Supports text, image, video, audio, and PDF input with a 1 million-token context window.

Capability profile

Category strength across 20 domains, via LLM Stats.

biology
90physics
90language
90chemistry
90multimodal
80vision
60reasoning
60math
50
general
50healthcare
50grounding
40factuality
40long context
40finance
30agents
10legal
0
Via LLM Stats

Benchmark breakdown

Independent evaluation scores, normalized to 0–100.

GPQA Diamond
82
Humanity’s Last Exam
16
MMLU-Pro
SciCode
42
LiveCodeBench
MATH-500
AIME 2025
τ²-Bench (agentic)
31
Terminal-Bench Hard
24
IFBench
77
Via Artificial Analysis

Providers

Inference hosts serving this model — their own pricing and measured performance, cheapest input first.

ProviderIn $/1MOut $/1MThroughputLatencyStatus
Google$0.25$1.50active
All providers

1 host · via LLM Stats

© 2026 NYSGPT2525 LLC