Nemotron 3 Ultra
TL;DR
Open-sourceOpen reasoning model for long-horizon autonomous agents.
Available on
Open-source models are available on every plan, including Go ($1/mo).
Switch with
Pick Nemotron 3 Ultra from the selector.
Intelligence Index
Nemotron 3 Ultra vs the top of the lineup
Claude Fable 5
GPT-5.6 Sol
Claude Opus 4.8
Nemotron 3 Ultra
Speed
~204
tokens / sec
Input
$0.60
per M tokens
Output
$2.40
per M tokens
Nemotron 3 Ultra in Command Code
Nemotron 3 Ultra is NVIDIA's open-source model — open reasoning model for long-horizon autonomous agents. It runs in Command Code with a 1M-token context window, switchable any time with /model.
It scores 37.8 on the Intelligence Index — #32 of the 38 scored models in the Command Code lineup.
Nemotron 3 Ultra specs at a glance
What Nemotron 3 Ultra accepts, how much it can hold in context, and what you need to run it.
| Spec | Nemotron 3 Ultra |
|---|---|
| Context window | 1M tokens |
| Input modalities | text |
| Reasoning | Yes |
| Release date | 2026-06-04 |
| Minimum plan | Go |
Nemotron 3 Ultra vs the Command Code lineup
Nemotron 3 Ultra alongside its nearest real alternatives in the lineup — every price straight from the billing tables.
| Model | Intelligence | Coding | Speed | Input $/M | Output $/M | Blended $/M | Context |
|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 59.9 | 76.5 | ~62 tok/s | $10.00 | $50.00 | $20 | 1M |
| Tencent Hy3 | 41.2 | 58.8 | ~58 tok/s | $0.00 | $0.00 | free | 262K |
| Qwen 3.7 Plus | 39 | 55.9 | ~52 tok/s | $0.40 | $1.60 | $0.70 | 1M |
| Kimi K2.5 | 38.1 | 46.8 | ~51 tok/s | $0.60 | $3.00 | $1.20 | 256K |
| MiniMax M2.7 | 38.1 | 52.6 | ~49 tok/s | $0.30 | $1.20 | $0.525 | 200K |
| Nemotron 3 Ultra (this page) | 37.8 | 49.3 | ~204 tok/s | $0.60 | $2.40 | $1.05 | 1M |
| MiMo V2.5 | 37.2 | 56.8 | ~88 tok/s | $0.14 | $0.28 | $0.175 | 1M |
What Nemotron 3 Ultra is best for
Nemotron 3 Ultra is a open-source model from NVIDIA, cheaper blended than most of the lineup ($1.05 per million tokens). It streams fast enough for interactive lookups.
The honest way to place it: run your own session with /model and compare against the lineup table above — the numbers on this page update as the registry and billing tables change.
When to switch away from Nemotron 3 Ultra
No single model wins every task. These are Nemotron 3 Ultra's computed nearest alternatives — one step up, one step down in cost, one for speed, one from the same family — each switchable mid-session with /model.
Switch to MiniMax M2.7
MiniMax M2.7 scores 38.1 on the Intelligence Index to Nemotron 3 Ultra's 37.8 — the nearest genuine step up — at $0.525 blended per million tokens versus $1.05. Reach for it when a task keeps hitting Nemotron 3 Ultra's ceiling.
Switch to MiniMax M3
MiniMax M3 runs about 62% cheaper blended ($0.3938 versus $1.05 per million tokens) while scoring 44.4 on the Intelligence Index. Switch down for high-volume work where Nemotron 3 Ultra's edge isn't earning its rate.
Switch to Gemini 3.1 Flash Lite
Gemini 3.1 Flash Lite streams ~300 tokens/sec to Nemotron 3 Ultra's ~204 at a comparable blended cost ($0.5625 per million tokens). Use it when iteration speed matters more than squeezing out the last point of quality.
What you pay for Nemotron 3 Ultra
Nemotron 3 Ultra is billed per token at the rates below — the same billing tables the Usage page charges against, so this page cannot quote a different price than you pay.
Blended cost (3:1 input:output, the shape of a typical coding session) works out to $1.05 per million tokens.
| Per 1M tokens | Input | Output | Cache read |
|---|---|---|---|
| All requests | $0.60 | $2.40 | $0.12 |
In Command Code: caching and taste-1
Open-source models are routed across multiple upstream providers for high availability. The price you see is the mean per-provider rate; the Usage page reflects what was actually charged.
Where supported by the upstream, prompt caching is on by default — cache reads are billed at $0.12 per million tokens versus $0.60 for fresh input.
taste-1 sits between the model and the agent loop, rewriting and reranking candidate edits to match your codebase conventions.
Plan availability
Nemotron 3 Ultra is an open-source model, available on every plan including Go.
Command Code is a subscription with model usage at API rates. Each plan ships with monthly LLM credits; credits roll over and never expire, and auto top-up keeps you running if you go over.
| Plan | Price/mo | LLM credits | Models |
|---|---|---|---|
| Go | $1 | $10 | Open-source only |
| Pro | $15 | $30 | Open-source + premium |
| Provider | $15 | Pay as you go | Open-source + premium |
| Max 10× | $100 | $150 | Open-source + premium |
| Max 20× | $200 | $300 | Open-source + premium |
| Teams Pro | $40 / seat | $40 / seat | Open-source + premium |
| Enterprise | Custom | Custom | Custom pool, SSO, audit logs |
Switching models with /model
In an interactive Command Code session, run /model to open the model selector. Pick Nemotron 3 Ultra and it applies to this session and to future sessions until you change it again. Premium models require Pro or higher; open-source models are available on every plan, including Go.
cmd # start an interactive session
/model # open the selector and pick Nemotron 3 UltraAll Command Code models, ranked by quality and speed
Quality is the Intelligence Index — an aggregate score across reasoning, math, coding, and knowledge evaluations. Speed is measured output tokens per second. Models without a published score are noted. This table is regenerated from the model registry, so it is always current.
| Model | Tier | Intelligence Index | Output speed |
|---|---|---|---|
| Claude Fable 5 | Premium | 59.9 | ~62 tok/s |
| GPT-5.6 Sol | Premium | 58.9 | ~77 tok/s |
| Claude Opus 4.8 | Premium | 55.7 | ~55 tok/s |
| GPT-5.6 Terra | Premium | 55 | ~155 tok/s |
| GPT-5.5 | Premium | 54.8 | ~81 tok/s |
| Grok 4.5 | Open-source | 53.8 | ~114 tok/s |
| Claude Opus 4.7 | Premium | 53.5 | ~52 tok/s |
| Claude Sonnet 5 | Premium | 53.4 | ~79 tok/s |
| GPT-5.4 | Premium | 51.4 | ~164 tok/s |
| GPT-5.6 Luna | Premium | 51.2 | ~234 tok/s |
| GLM-5.2 | Open-source | 51.1 | ~208 tok/s |
| Muse Spark 1.1 | Premium | 50.6 | ~130 tok/s |
| Gemini 3.5 Flash | Premium | 50.2 | ~236 tok/s |
| Claude Sonnet 4.6 | Premium | 47.2 | ~55 tok/s |
| Qwen 3.7 Max | Open-source | 46 | ~196 tok/s |
| MiniMax M3 | Open-source | 44.4 | ~113 tok/s |
| DeepSeek V4 Pro | Open-source | 44.3 | ~62 tok/s |
| GPT-5.3 Codex | Premium | 44.3 | ~106 tok/s |
| Kimi K2.6 | Open-source | 44.2 | ~43 tok/s |
| MiMo V2.5 Pro | Open-source | 42.2 | ~56 tok/s |
| Kimi K2.7 Code | Open-source | 41.9 | ~46 tok/s |
| Tencent Hy3 | Open-source | 41.2 | ~58 tok/s |
| DeepSeek V4 Flash | Open-source | 40.3 | ~106 tok/s |
| GLM-5.1 | Open-source | 40.2 | ~81 tok/s |
| Qwen 3.6 Max Preview | Open-source | 40 | ~46 tok/s |
| GPT-5.4 Mini | Premium | 40 | ~171 tok/s |
| Qwen 3.6 Plus | Open-source | 39.6 | ~53 tok/s |
| GLM-5 | Open-source | 39.5 | ~51 tok/s |
| Qwen 3.7 Plus | Open-source | 39 | ~52 tok/s |
| Kimi K2.5 | Open-source | 38.1 | ~51 tok/s |
| MiniMax M2.7 | Open-source | 38.1 | ~49 tok/s |
| Nemotron 3 Ultra (this page) | Open-source | 37.8 | ~204 tok/s |
| MiMo V2.5 | Open-source | 37.2 | ~88 tok/s |
| MiniMax M2.5 | Open-source | 33.7 | ~78 tok/s |
| Step 3.7 Flash | Open-source | 30.3 | ~407 tok/s |
| Step 3.5 Flash | Open-source | 26 | ~207 tok/s |
| Gemini 3.1 Flash Lite | Premium | 25 | ~300 tok/s |
| Claude Haiku 4.5 | Premium | 23.7 | ~103 tok/s |
| Kimi K3 | Open-source | Not yet scored | — |
| Kimi K2.7 Code HighSpeed | Open-source | Not yet scored | — |
| GLM-5.2 Fast | Open-source | Not yet scored | — |
| Inkling | Open-source | Not yet scored | — |
| Fugu Ultra | Premium | Not yet scored | — |
Frequently asked questions
Nemotron 3 Ultra or MiniMax M2.7?
MiniMax M2.7 scores higher on the Intelligence Index (38.1 vs 37.8) at $0.525 blended per million tokens against Nemotron 3 Ultra's $1.05. Default to Nemotron 3 Ultra and switch up when a task keeps stalling.
Nemotron 3 Ultra or MiniMax M3?
MiniMax M3 is about 62% cheaper blended ($0.3938 vs $1.05 per million tokens), scoring 44.4 on the Intelligence Index. Use MiniMax M3 for volume work and Nemotron 3 Ultra where its edge earns the difference.
How much does Nemotron 3 Ultra cost in Command Code?
$0.60 per million input tokens and $2.40 per million output tokens, with cache reads at $0.12. In an agent loop, cached context brings effective input to roughly $0.264 per million tokens.
What plan do I need for Nemotron 3 Ultra?
Nemotron 3 Ultra is available on every plan, including Go at $1/mo.
How good is Nemotron 3 Ultra at coding?
Nemotron 3 Ultra scores 49.3 on the Coding Index — the Artificial Analysis sub-score most predictive of coding-agent performance — alongside 37.8 overall.
Does Nemotron 3 Ultra support image input and reasoning?
Nemotron 3 Ultra is text-only; paste code and logs rather than screenshots. It supports reasoning.
Which Command Code model should I use?
Claude Fable 5 currently leads the lineup on the Intelligence Index (59.9). Grok 4.5 (53.8) leads the open-weights tier, available on every plan. For fast lookups, Step 3.7 Flash streams ~407 tok/s. There is no single right answer — switch per session with /model and let the task pick the model.
Can I mix Nemotron 3 Ultra with other models in a workflow?
Yes. Switch per session using /model. Common pattern: keep a default model and switch up for hard problems or down for quick lookups as the task calls for it.
Are open-source model prices fixed?
Open-source models are routed across multiple upstream providers for high availability. The price listed for each is the mean per-provider rate. Actual cost on a given request may vary slightly. The Usage page reflects the price charged.
What does "550b-a55b" mean?
It is a 550-billion-parameter mixture-of-experts model with 55 billion active parameters per token — frontier-class capacity at a lower serving cost, reflected in the price.
Why no Intelligence Index for Nemotron 3 Ultra?
Public aggregate benchmarks have not yet been published for Nemotron 3 Ultra in the current Intelligence Index format. The model is available and routed normally.
Is Command Code free to try?
The Go plan starts at $1/mo with $10 in LLM credits. It covers open-source models only. Pro at $15/mo unlocks premium models with $30 in LLM credits.
Does Command Code train on my code?
No. Command Code does not train on your code or store your code snippets. taste-1 data is stored locally in your project directory.
Where can I track my usage?
The Usage page in Studio shows per-request cost, token counts, and which model ran. Settings > Billing lets you change plans, buy credits, or enable auto top-up.
Does Command Code replace my editor?
No. Command Code is editor-agnostic — it runs as a CLI and works alongside any editor (Cursor, VS Code, Zed, JetBrains, Neovim, etc.).
Related reading
- Kimi K2.5 in Command Codemultimodal frontend coding
- MiniMax M2.7 in Command Codeend-to-end software engineering agent
- Every model, one referenceThe docs list of all Command Code models with ids and context windows.
- Pricing, limits, and dealsThe canonical price table, running deals, and usage estimates.
Ship code that matches your taste
Command Code is the AI coding agent that continuously learns your taste. Start for $1.