Gemini 3.1 Flash Lite

google/gemini-3.1-flash-lite

high-volume workhorse model with implicit caching.

cmd --model google/gemini-3.1-flash-lite
Intelligence index
25.6
Output speed
not yet scored
Input
$0.25 /M
Output
$1.50 /M
Cache read
$0.03 /M
Agent-loop cost
$0.10 /M in
Context window
1M tokens
Released
March 3, 2026
Modalities

vs. the lineup

Gemini 3.1 Flash Lite beside its stablemates and nearest rivals. The ◆ marks the best value in each column across every row shown.

pin a rival:
ModelIntelligenceCodingSpeedInput $/MOutput $/MBlended $/MContext
Muse Spark 1.2 Contributor56.872.2$0.10$0.20$0.131.05M
Gemini 3.5 Flash5270.1$1.50$9$3.381M
Gemini 3.6 Flash51.669.2213.7$1.50$7.50$31M
Gemini 3.5 Flash Lite37.449.3342.7$0.30$2.50$0.851M
Gemini 3.1 Flash Lite25.634.7$0.25$1.50$0.561M
Gemini 3.7 Flash$0.75$3.75$1.501.05M
DeepSeek V4 Pro (latest)pinned45.359.463.3$0.66$1.98$0.991M
DeepSeek V4 Flash (latest)pinned5269.1115.9$0.22$0.66$0.331M

coding performance

The Intelligence Index and its sub-scores, ranked against every scored model in the catalog. A metric that has not been measured for Gemini 3.1 Flash Lite has been left empty.

Coding Index
34.7
#40 of 40 scored
Terminal-Bench
31.1
#40 of 40 scored
Intelligence Index
25.6
#45 of 46 scored
Long-context reasoning
71.3
reasoning across a long context
SciCode
41.9
scientific coding
GPQA Diamond
82.2
graduate-level QA

usage calculator

How far a month of credits goes on Gemini 3.1 Flash Lite.

Gemini 3.1 Flash Lite runs on Pro and up — from $20/mo.
Input tokensfresh prompt
800
Output tokensmodel reply
180
Cache read tokensre-read context
50K
cost / request $0.0020 · in $0.25 · out $1.50 · cache $0.03 per M
fresh input 10%output 14%cache reads 76%
Requests / 30 days
10K
$20 credits ÷ $0.0020 per request
~2.0K quick fixes~406 bug fixes~68 feature PRs

what real work costs

Real coding tasks priced end to end on Gemini 3.1 Flash Lite, from a quick lookup to a full-repo agent run.

One agent task
$0.03
180K in at 75% cache hit, 12K out
What you pay
$0.17 /M
all-in across every token that task touched
Sticker input
$0.25 /M
cache reads bill at $0.03 /M instead
TaskTokens in · outGemini 3.1 Flash LiteMuse Spark 1.2 ContributorClaude Haiku 4.5
Quick lookup / one-liner8K · 1K$0.0019$0.0003$0.0065
Review a 500-line PR60K · 4K$0.01$0.0027$0.04
Fix a bug (agent loop)180K · 12K$0.03$0.0072$0.12
Refactor a module320K · 20K$0.06$0.01$0.21
Full-repo agent run900K · 45K$0.14$0.03$0.49

frequently asked

What is Gemini 3.1 Flash Lite best for?
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic workflows, simple data extraction, and applications where responsiveness and API cost are the primary constraints. Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash.
How much does Gemini 3.1 Flash Lite cost?
$0.25/M input and $1.50/M output, cache reads $0.03/M. In an agent loop most input is cache-read, so the effective input rate is about $0.10/M.
Which plan do I need?
Available on Pro and above.
How do I switch to it?
Run cmd --model google/gemini-3.1-flash-lite, or type /model in a session and pick it. You can switch mid-session without losing context.

Ship code that matches your taste

Command Code is the AI coding agent that continuously learns your taste. Start for $1.

Benchmarks from Artificial Analysis (v4.1)commandcode.ai/models/gemini-3-1-flash-lite