← All models
Balanced — strong quality at a mid priceOpen Weights
Nemotron 3 Ultra Pricing
NVIDIA · nvidia-nemotron-3-ultra-550b-a55b
Input / 1M tokens
$0.600
Output / 1M tokens
$2.40
Context window
262,144 tokens
≈ 350 pages of text
Max output
182,520 tokens
What it costs in practice
Typical request (1,200 in + 400 out tokens)$0.0017
1,000 requests / month$1.68/mo10,000 requests / month$16.80/mo100,000 requests / month$168.00/moEstimate your own workloadOpens the calculator with Nemotron 3 Ultra preloaded — adjust volume and token counts there.
Price history
List price per 1M tokens since we started tracking in Jun 2026. Each step marks a price change — hover a point for the exact date and rate.
What it excels at
Strong performance on long-context tasks enabled by the 1M token window and 16k output limit.
The business tradeoff
Open-weights deployment requires substantial self-hosted compute and introduces variance in latency and throughput.
Cheaper from NVIDIA
Head-to-head
Vendor list rates, as of Sep 14, 2026 · source: openrouter · per-request examples assume 1,200 input + 400 output tokens.