Models

Nemotron 3 Ultra

TL;DR

Open-source

Open reasoning model for long-horizon autonomous agents.

Available on

GoProProviderMax 10×Max 20×Teams ProEnterprise

Open-source models are available on every plan, including Go ($1/mo).

Switch with

/model

Pick Nemotron 3 Ultra from the selector.

Intelligence Index

Nemotron 3 Ultra vs the top of the lineup

37.8

Claude Fable 5

59.9

GPT-5.6 Sol

58.9

Claude Opus 4.8

55.7

Nemotron 3 Ultra

37.8

Speed

~204

tokens / sec

Input

$0.60

per M tokens

Output

$2.40

per M tokens

Nemotron 3 Ultra in Command Code

Nemotron 3 Ultra is NVIDIA's open-source model — open reasoning model for long-horizon autonomous agents. It runs in Command Code with a 1M-token context window, switchable any time with /model.

It scores 37.8 on the Intelligence Index — #32 of the 38 scored models in the Command Code lineup.

Nemotron 3 Ultra specs at a glance

What Nemotron 3 Ultra accepts, how much it can hold in context, and what you need to run it.

SpecNemotron 3 Ultra
Context window1M tokens
Input modalitiestext
ReasoningYes
Release date2026-06-04
Minimum planGo

Nemotron 3 Ultra vs the Command Code lineup

Nemotron 3 Ultra alongside its nearest real alternatives in the lineup — every price straight from the billing tables.

ModelIntelligenceCodingSpeedInput $/MOutput $/MBlended $/MContext
Claude Fable 559.976.5~62 tok/s$10.00$50.00$201M
Tencent Hy341.258.8~58 tok/s$0.00$0.00free262K
Qwen 3.7 Plus3955.9~52 tok/s$0.40$1.60$0.701M
Kimi K2.538.146.8~51 tok/s$0.60$3.00$1.20256K
MiniMax M2.738.152.6~49 tok/s$0.30$1.20$0.525200K
Nemotron 3 Ultra (this page)37.849.3~204 tok/s$0.60$2.40$1.051M
MiMo V2.537.256.8~88 tok/s$0.14$0.28$0.1751M

What Nemotron 3 Ultra is best for

Nemotron 3 Ultra is a open-source model from NVIDIA, cheaper blended than most of the lineup ($1.05 per million tokens). It streams fast enough for interactive lookups.

The honest way to place it: run your own session with /model and compare against the lineup table above — the numbers on this page update as the registry and billing tables change.

When to switch away from Nemotron 3 Ultra

No single model wins every task. These are Nemotron 3 Ultra's computed nearest alternatives — one step up, one step down in cost, one for speed, one from the same family — each switchable mid-session with /model.

Switch to MiniMax M2.7

MiniMax M2.7 scores 38.1 on the Intelligence Index to Nemotron 3 Ultra's 37.8 — the nearest genuine step up — at $0.525 blended per million tokens versus $1.05. Reach for it when a task keeps hitting Nemotron 3 Ultra's ceiling.

Switch to MiniMax M3

MiniMax M3 runs about 62% cheaper blended ($0.3938 versus $1.05 per million tokens) while scoring 44.4 on the Intelligence Index. Switch down for high-volume work where Nemotron 3 Ultra's edge isn't earning its rate.

Switch to Gemini 3.1 Flash Lite

Gemini 3.1 Flash Lite streams ~300 tokens/sec to Nemotron 3 Ultra's ~204 at a comparable blended cost ($0.5625 per million tokens). Use it when iteration speed matters more than squeezing out the last point of quality.

What you pay for Nemotron 3 Ultra

Nemotron 3 Ultra is billed per token at the rates below — the same billing tables the Usage page charges against, so this page cannot quote a different price than you pay.

Blended cost (3:1 input:output, the shape of a typical coding session) works out to $1.05 per million tokens.

Per 1M tokensInputOutputCache read
All requests$0.60$2.40$0.12

In Command Code: caching and taste-1

Open-source models are routed across multiple upstream providers for high availability. The price you see is the mean per-provider rate; the Usage page reflects what was actually charged.

Where supported by the upstream, prompt caching is on by default — cache reads are billed at $0.12 per million tokens versus $0.60 for fresh input.

taste-1 sits between the model and the agent loop, rewriting and reranking candidate edits to match your codebase conventions.

Plan availability

Nemotron 3 Ultra is an open-source model, available on every plan including Go.

Command Code is a subscription with model usage at API rates. Each plan ships with monthly LLM credits; credits roll over and never expire, and auto top-up keeps you running if you go over.

PlanPrice/moLLM creditsModels
Go$1$10Open-source only
Pro$15$30Open-source + premium
Provider$15Pay as you goOpen-source + premium
Max 10×$100$150Open-source + premium
Max 20×$200$300Open-source + premium
Teams Pro$40 / seat$40 / seatOpen-source + premium
EnterpriseCustomCustomCustom pool, SSO, audit logs

Switching models with /model

In an interactive Command Code session, run /model to open the model selector. Pick Nemotron 3 Ultra and it applies to this session and to future sessions until you change it again. Premium models require Pro or higher; open-source models are available on every plan, including Go.

cmd               # start an interactive session
/model            # open the selector and pick Nemotron 3 Ultra

All Command Code models, ranked by quality and speed

Quality is the Intelligence Index — an aggregate score across reasoning, math, coding, and knowledge evaluations. Speed is measured output tokens per second. Models without a published score are noted. This table is regenerated from the model registry, so it is always current.

ModelTierIntelligence IndexOutput speed
Claude Fable 5Premium59.9~62 tok/s
GPT-5.6 SolPremium58.9~77 tok/s
Claude Opus 4.8Premium55.7~55 tok/s
GPT-5.6 TerraPremium55~155 tok/s
GPT-5.5Premium54.8~81 tok/s
Grok 4.5Open-source53.8~114 tok/s
Claude Opus 4.7Premium53.5~52 tok/s
Claude Sonnet 5Premium53.4~79 tok/s
GPT-5.4Premium51.4~164 tok/s
GPT-5.6 LunaPremium51.2~234 tok/s
GLM-5.2Open-source51.1~208 tok/s
Muse Spark 1.1Premium50.6~130 tok/s
Gemini 3.5 FlashPremium50.2~236 tok/s
Claude Sonnet 4.6Premium47.2~55 tok/s
Qwen 3.7 MaxOpen-source46~196 tok/s
MiniMax M3Open-source44.4~113 tok/s
DeepSeek V4 ProOpen-source44.3~62 tok/s
GPT-5.3 CodexPremium44.3~106 tok/s
Kimi K2.6Open-source44.2~43 tok/s
MiMo V2.5 ProOpen-source42.2~56 tok/s
Kimi K2.7 CodeOpen-source41.9~46 tok/s
Tencent Hy3Open-source41.2~58 tok/s
DeepSeek V4 FlashOpen-source40.3~106 tok/s
GLM-5.1Open-source40.2~81 tok/s
Qwen 3.6 Max PreviewOpen-source40~46 tok/s
GPT-5.4 MiniPremium40~171 tok/s
Qwen 3.6 PlusOpen-source39.6~53 tok/s
GLM-5Open-source39.5~51 tok/s
Qwen 3.7 PlusOpen-source39~52 tok/s
Kimi K2.5Open-source38.1~51 tok/s
MiniMax M2.7Open-source38.1~49 tok/s
Nemotron 3 Ultra (this page)Open-source37.8~204 tok/s
MiMo V2.5Open-source37.2~88 tok/s
MiniMax M2.5Open-source33.7~78 tok/s
Step 3.7 FlashOpen-source30.3~407 tok/s
Step 3.5 FlashOpen-source26~207 tok/s
Gemini 3.1 Flash LitePremium25~300 tok/s
Claude Haiku 4.5Premium23.7~103 tok/s
Kimi K3Open-sourceNot yet scored
Kimi K2.7 Code HighSpeedOpen-sourceNot yet scored
GLM-5.2 FastOpen-sourceNot yet scored
InklingOpen-sourceNot yet scored
Fugu UltraPremiumNot yet scored

Frequently asked questions

Nemotron 3 Ultra or MiniMax M2.7?

MiniMax M2.7 scores higher on the Intelligence Index (38.1 vs 37.8) at $0.525 blended per million tokens against Nemotron 3 Ultra's $1.05. Default to Nemotron 3 Ultra and switch up when a task keeps stalling.

Nemotron 3 Ultra or MiniMax M3?

MiniMax M3 is about 62% cheaper blended ($0.3938 vs $1.05 per million tokens), scoring 44.4 on the Intelligence Index. Use MiniMax M3 for volume work and Nemotron 3 Ultra where its edge earns the difference.

How much does Nemotron 3 Ultra cost in Command Code?

$0.60 per million input tokens and $2.40 per million output tokens, with cache reads at $0.12. In an agent loop, cached context brings effective input to roughly $0.264 per million tokens.

What plan do I need for Nemotron 3 Ultra?

Nemotron 3 Ultra is available on every plan, including Go at $1/mo.

How good is Nemotron 3 Ultra at coding?

Nemotron 3 Ultra scores 49.3 on the Coding Index — the Artificial Analysis sub-score most predictive of coding-agent performance — alongside 37.8 overall.

Does Nemotron 3 Ultra support image input and reasoning?

Nemotron 3 Ultra is text-only; paste code and logs rather than screenshots. It supports reasoning.

Which Command Code model should I use?

Claude Fable 5 currently leads the lineup on the Intelligence Index (59.9). Grok 4.5 (53.8) leads the open-weights tier, available on every plan. For fast lookups, Step 3.7 Flash streams ~407 tok/s. There is no single right answer — switch per session with /model and let the task pick the model.

Can I mix Nemotron 3 Ultra with other models in a workflow?

Yes. Switch per session using /model. Common pattern: keep a default model and switch up for hard problems or down for quick lookups as the task calls for it.

Are open-source model prices fixed?

Open-source models are routed across multiple upstream providers for high availability. The price listed for each is the mean per-provider rate. Actual cost on a given request may vary slightly. The Usage page reflects the price charged.

What does "550b-a55b" mean?

It is a 550-billion-parameter mixture-of-experts model with 55 billion active parameters per token — frontier-class capacity at a lower serving cost, reflected in the price.

Why no Intelligence Index for Nemotron 3 Ultra?

Public aggregate benchmarks have not yet been published for Nemotron 3 Ultra in the current Intelligence Index format. The model is available and routed normally.

Is Command Code free to try?

The Go plan starts at $1/mo with $10 in LLM credits. It covers open-source models only. Pro at $15/mo unlocks premium models with $30 in LLM credits.

Does Command Code train on my code?

No. Command Code does not train on your code or store your code snippets. taste-1 data is stored locally in your project directory.

Where can I track my usage?

The Usage page in Studio shows per-request cost, token counts, and which model ran. Settings > Billing lets you change plans, buy credits, or enable auto top-up.

Does Command Code replace my editor?

No. Command Code is editor-agnostic — it runs as a CLI and works alongside any editor (Cursor, VS Code, Zed, JetBrains, Neovim, etc.).

Related reading

Ship code that matches your taste

Command Code is the AI coding agent that continuously learns your taste. Start for $1.