DeepSeek V4 Flash

deepseek/deepseek-v4-flash

fast hybrid-attention reasoning.

cmd --model deepseek/deepseek-v4-flash
Intelligence index
40.3
Output speed
122 tok/s
Input
$0.14 /M
Output
$0.28 /M
Cache read
$0.00 /M
Agent-loop cost
$0.04 /M in
Context window
1M tokens
Released
April 24, 2026
Modalities

vs. the lineup

DeepSeek V4 Flash beside its stablemates and nearest rivals. The ◆ marks the best value in each column across every row shown.

pin a rival:
ModelIntelligenceCodingSpeedInput $/MOutput $/MBlended $/MContext
DeepSeek V4 Pro44.359.470.9$0.43$0.87$0.541M
Tencent Hy341.258.865.2$0.14$0.58$0.25262K
Inkling40.752.156.5$1$4.05$1.76256K
DeepSeek V4 Flash40.356.2122$0.14$0.28$0.181M
GLM-5.140.255.868.7$1.40$4.40$2.15200K
Step 3.7 Flash30.339.6399.5$0.20$1.15$0.44256K
Kimi K3pinned57.176.234.5$3$15$61M
Kimi K2.7 Codepinned41.960.844.4$0.95$4$1.71256K

coding performance

The Intelligence Index and its sub-scores, ranked against every scored model in the catalog. A metric that has not been measured for DeepSeek V4 Flash has been left empty.

Coding Index
56.2
#24 of 34 scored
Terminal-Bench
61.8
#25 of 34 scored
Intelligence Index
40.3
#25 of 40 scored
Long-context reasoning
63
reasoning across a long context
SciCode
44.9
scientific coding
GPQA Diamond
89.4
graduate-level QA

usage calculator

How far a month of credits goes on DeepSeek V4 Flash.

Input tokensfresh prompt
800
Output tokensmodel reply
180
Cache read tokensre-read context
50K
cost / request $0.0003 · in $0.14 · out $0.28 · cache $0.00 per M
fresh input 37%output 17%cache reads 46%
Requests / 30 days
33K
$10 credits ÷ $0.0003 per request
~6.6K quick fixes~1.3K bug fixes~220 feature PRs

what real work costs

Real coding tasks priced end to end on DeepSeek V4 Flash, from a quick lookup to a full-repo agent run.

One agent task
$0.01
180K in at 75% cache hit, 12K out
What you pay
$0.05 /M
all-in across every token that task touched
Sticker input
$0.14 /M
cache reads bill at $0.00 /M instead
TaskTokens in · outDeepSeek V4 FlashClaude Haiku 4.5
Quick lookup / one-liner8K · 1K$0.0004$0.0065
Review a 500-line PR60K · 4K$0.0038$0.04
Fix a bug (agent loop)180K · 12K$0.01$0.12
Refactor a module320K · 20K$0.02$0.21
Full-repo agent run900K · 45K$0.04$0.49

frequently asked

What is DeepSeek V4 Flash best for?+
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance. The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.
How much does DeepSeek V4 Flash cost?+
$0.14/M input and $0.28/M output, cache reads $0.00/M. In an agent loop most input is cache-read, so the effective input rate is about $0.04/M.
Which plan do I need?+
Available on Go and above.
How do I switch to it?+
Run cmd --model deepseek/deepseek-v4-flash, or type /model in a session and pick it. You can switch mid-session without losing context.

Ship code that matches your taste

Command Code is the AI coding agent that continuously learns your taste. Start for $1.

commandcode.ai/models/deepseek-v4-flash