← All models
Flash — fastest and cheapest, lighter reasoningOpen Weights
Nemotron 3 Nano 30B A3B Pricing
NVIDIA · nvidia-nemotron-3-nano-30b-a3b
Input / 1M tokens
$0.050
Output / 1M tokens
$0.200
Context window
262,144 tokens
≈ 350 pages of text
Max output
228k tokens
What it costs in practice
Typical request (1,200 in + 400 out tokens)$0.0001
1,000 requests / month$0.14/mo10,000 requests / month$1.40/mo100,000 requests / month$14.00/moEstimate your own workloadOpens the calculator with Nemotron 3 Nano 30B A3B preloaded — adjust volume and token counts there.
Price history
List price per 1M tokens since we started tracking in Jun 2026. The lines are flat because the price hasn't changed since then.
What it excels at
Handles long-context workloads up to 262k tokens at low per-token cost; suitable for high-volume inference when self-hosted.
The business tradeoff
30B-scale reasoning trails larger models on complex tasks; open-weights deployment shifts all latency, throughput, and reliability costs to the operator.
Head-to-head
Vendor list rates, as of Jul 28, 2026 · source: openrouter · per-request examples assume 1,200 input + 400 output tokens.