← All models
Balanced — strong quality at a mid priceOpen Weights
Nemotron 3 Ultra Pricing
NVIDIA · nvidia-nemotron-3-ultra-550b-a55b
Input / 1M tokens
$0.600
Output / 1M tokens
$3.60
Context window
512,288 tokens
≈ 683 pages of text
Max output
—
What it costs in practice
Typical request (1,200 in + 400 out tokens)$0.0022
1,000 requests / month$2.16/mo10,000 requests / month$21.60/mo100,000 requests / month$216.00/moEstimate your own workloadOpens the calculator with Nemotron 3 Ultra preloaded — adjust volume and token counts there.
Price history
List price per 1M tokens since we started tracking in Jun 2026. Each step marks a price change — hover a point for the exact date and rate.
What it excels at
Strong performance on long-context tasks enabled by the 1M token window and 16k output limit.
The business tradeoff
Open-weights deployment requires substantial self-hosted compute and introduces variance in latency and throughput.
Cheaper from NVIDIA
Head-to-head
Vendor list rates, as of Jul 30, 2026 · source: openrouter · per-request examples assume 1,200 input + 400 output tokens.