Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b
open reasoning model for long-horizon autonomous agents.
cmd --model nvidia/nemotron-3-ultra-550b-a55b
Intelligence index
22.9
Coding index
49.3
Input
$0.60 /M
Output
$2.40 /M
Cache read
$0.12 /M
Agent-loop cost
$0.26 /M in
Context window
1M tokens
Released
June 4, 2026
Modalities
→
vs. the lineup
Nemotron 3 Ultra beside its stablemates and nearest rivals. The ◆ marks the best value in each column across every row shown.
pin a rival:
| Model | Intelligence | Coding | Input $/M | Output $/M | Blended $/M | Context |
|---|---|---|---|---|---|---|
| MiMo V2.6 Pro | 46.3◆ | —◆ | $0.43◆ | $0.87◆ | $0.54◆ | 1.05M◆ |
| GPT-5.4 Mini | 24.1◆ | 56.1◆ | $0.75◆ | $4.50◆ | $1.69◆ | 400K◆ |
| Kimi K2.5 | 23.5◆ | 46.8◆ | $0.60◆ | $3◆ | $1.20◆ | 256K◆ |
| Nemotron 3 Ultra ◆ | 22.9◆ | 49.3◆ | $0.60◆ | $2.40◆ | $1.05◆ | 1M◆ |
| MiniMax M2.7 | 22.8◆ | 52.6◆ | $0.30◆ | $1.20◆ | $0.52◆ | 200K◆ |
| MiniMax M2.5 | 22.8◆ | —◆ | $0.30◆ | $1.20◆ | $0.52◆ | 200K◆ |
| DeepSeek V4 Pro (latest)pinned | 36◆ | 68.8◆ | $0.66◆ | $1.98◆ | $0.99◆ | 1M◆ |
| DeepSeek V4 Flash (latest)pinned | 34◆ | 69.1◆ | $0.15◆ | $0.60◆ | $0.26◆ | 1M◆ |
coding performance
The Intelligence Index and its sub-scores, ranked against every scored model in the catalog. A metric that has not been measured for Nemotron 3 Ultra has been left empty.
Coding Index
49.3
#47 of 52 scored
Terminal-Bench
53.9
#47 of 52 scored
Intelligence Index
22.9
#60 of 68 scored
Long-context reasoning
79.3
reasoning across a long context
SciCode
40.3
scientific coding
GPQA Diamond
86.7
graduate-level QA
usage calculator
How far a month of credits goes on Nemotron 3 Ultra.
Input tokensfresh prompt
800Output tokensmodel reply
180Cache read tokensre-read context
50Kcost / request $0.0069 · in $0.60 · out $2.40 · cache $0.12 per M
fresh input 7%output 6%cache reads 87%
Requests / 30 days
868
$6 credits ÷ $0.0069 per request
~174 quick fixes~35 bug fixes~6 feature PRs
what real work costs
Real coding tasks priced end to end on Nemotron 3 Ultra, from a quick lookup to a full-repo agent run.
One agent task
$0.07
180K in at 75% cache hit, 12K out
What you pay
$0.38 /M
all-in across every token that task touched
Sticker input
$0.60 /M
cache reads bill at $0.12 /M instead
| Task | Tokens in · out | Nemotron 3 Ultra | MiMo V2.6 Pro | Claude Haiku 4.5 |
|---|---|---|---|---|
| Quick lookup / one-liner | 8K · 1K | $0.0037 | $0.0012 | $0.0065 |
| Review a 500-line PR | 60K · 4K | $0.03 | $0.01 | $0.04 |
| Fix a bug (agent loop) | 180K · 12K | $0.07 | $0.03 | $0.12 |
| Refactor a module | 320K · 20K | $0.13 | $0.06 | $0.21 |
| Full-repo agent run | 900K · 45K | $0.31 | $0.13 | $0.49 |
frequently asked
How much does Nemotron 3 Ultra cost?
$0.60/M input and $2.40/M output, cache reads $0.12/M. In an agent loop most input is cache-read, so the effective input rate is about $0.26/M.Which plan do I need?
Available on Go and above.
How do I switch to it?
Run
cmd --model nvidia/nemotron-3-ultra-550b-a55b, or type /model in a session and pick it. You can switch mid-session without losing context.Ship code that matches your taste
Command Code is the AI coding agent that continuously learns your taste. Start for $1.
Benchmarks from Artificial Analysis (v4.3)commandcode.ai/models/nemotron-3-ultra-550b-a55b