Gemini 3.1 Flash Lite
google/gemini-3.1-flash-lite
high-volume workhorse model with implicit caching.
cmd --model google/gemini-3.1-flash-lite
Intelligence index
25.6
Output speed
not yet scored
Input
$0.25 /M
Output
$1.50 /M
Cache read
$0.03 /M
Agent-loop cost
$0.10 /M in
Context window
1M tokens
Released
March 3, 2026
Modalities
→
vs. the lineup
Gemini 3.1 Flash Lite beside its stablemates and nearest rivals. The ◆ marks the best value in each column across every row shown.
pin a rival:
| Model | Intelligence | Coding | Speed | Input $/M | Output $/M | Blended $/M | Context |
|---|---|---|---|---|---|---|---|
| Muse Spark 1.2 Contributor | 56.8◆ | 72.2◆ | —◆ | $0.10◆ | $0.20◆ | $0.13◆ | 1.05M◆ |
| Gemini 3.5 Flash | 52◆ | 70.1◆ | —◆ | $1.50◆ | $9◆ | $3.38◆ | 1M◆ |
| Gemini 3.6 Flash | 51.6◆ | 69.2◆ | 213.7◆ | $1.50◆ | $7.50◆ | $3◆ | 1M◆ |
| Gemini 3.5 Flash Lite | 37.4◆ | 49.3◆ | 342.7◆ | $0.30◆ | $2.50◆ | $0.85◆ | 1M◆ |
| Gemini 3.1 Flash Lite ◆ | 25.6◆ | 34.7◆ | —◆ | $0.25◆ | $1.50◆ | $0.56◆ | 1M◆ |
| Gemini 3.7 Flash | —◆ | —◆ | —◆ | $0.75◆ | $3.75◆ | $1.50◆ | 1.05M◆ |
| DeepSeek V4 Pro (latest)pinned | 45.3◆ | 59.4◆ | 63.3◆ | $0.66◆ | $1.98◆ | $0.99◆ | 1M◆ |
| DeepSeek V4 Flash (latest)pinned | 52◆ | 69.1◆ | 115.9◆ | $0.22◆ | $0.66◆ | $0.33◆ | 1M◆ |
coding performance
The Intelligence Index and its sub-scores, ranked against every scored model in the catalog. A metric that has not been measured for Gemini 3.1 Flash Lite has been left empty.
Coding Index
34.7
#40 of 40 scored
Terminal-Bench
31.1
#40 of 40 scored
Intelligence Index
25.6
#45 of 46 scored
Long-context reasoning
71.3
reasoning across a long context
SciCode
41.9
scientific coding
GPQA Diamond
82.2
graduate-level QA
usage calculator
How far a month of credits goes on Gemini 3.1 Flash Lite.
Gemini 3.1 Flash Lite runs on Pro and up — from $20/mo.
Input tokensfresh prompt
800Output tokensmodel reply
180Cache read tokensre-read context
50Kcost / request $0.0020 · in $0.25 · out $1.50 · cache $0.03 per M
fresh input 10%output 14%cache reads 76%
Requests / 30 days
10K
$20 credits ÷ $0.0020 per request
~2.0K quick fixes~406 bug fixes~68 feature PRs
what real work costs
Real coding tasks priced end to end on Gemini 3.1 Flash Lite, from a quick lookup to a full-repo agent run.
One agent task
$0.03
180K in at 75% cache hit, 12K out
What you pay
$0.17 /M
all-in across every token that task touched
Sticker input
$0.25 /M
cache reads bill at $0.03 /M instead
| Task | Tokens in · out | Gemini 3.1 Flash Lite | Muse Spark 1.2 Contributor | Claude Haiku 4.5 |
|---|---|---|---|---|
| Quick lookup / one-liner | 8K · 1K | $0.0019 | $0.0003 | $0.0065 |
| Review a 500-line PR | 60K · 4K | $0.01 | $0.0027 | $0.04 |
| Fix a bug (agent loop) | 180K · 12K | $0.03 | $0.0072 | $0.12 |
| Refactor a module | 320K · 20K | $0.06 | $0.01 | $0.21 |
| Full-repo agent run | 900K · 45K | $0.14 | $0.03 | $0.49 |
frequently asked
What is Gemini 3.1 Flash Lite best for?
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic workflows, simple data extraction, and applications where responsiveness and API cost are the primary constraints. Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash.
How much does Gemini 3.1 Flash Lite cost?
$0.25/M input and $1.50/M output, cache reads $0.03/M. In an agent loop most input is cache-read, so the effective input rate is about $0.10/M.Which plan do I need?
Available on Pro and above.
How do I switch to it?
Run
cmd --model google/gemini-3.1-flash-lite, or type /model in a session and pick it. You can switch mid-session without losing context.Ship code that matches your taste
Command Code is the AI coding agent that continuously learns your taste. Start for $1.
Benchmarks from Artificial Analysis (v4.1)commandcode.ai/models/gemini-3-1-flash-lite