← All models
Balanced — strong quality at a mid priceOpen Weights
Hermes 3 405B Instruct Pricing
Nous · nousresearch-hermes-3-llama-3-1-405b
Input / 1M tokens
$1.00
Output / 1M tokens
$1.00
Context window
131,072 tokens
≈ 175 pages of text
Max output
16,384 tokens
What it costs in practice
Typical request (1,200 in + 400 out tokens)$0.0016
1,000 requests / month$1.60/mo10,000 requests / month$16.00/mo100,000 requests / month$160.00/moEstimate your own workloadOpens the calculator with Hermes 3 405B Instruct preloaded — adjust volume and token counts there.
Price history
Only one price point recorded so far.
The staircase chart appears once a price change is detected.
What it excels at
Strong instruction following and tool-use performance on long contexts up to 128k tokens. Competitive reasoning quality at very low per-token cost.
The business tradeoff
Output quality and latency depend heavily on the inference provider; self-hosting a 405B model requires substantial GPU infrastructure.
Cheaper from Nous
Head-to-head
Vendor list rates, as of Jun 15, 2026 · source: openrouter · per-request examples assume 1,200 input + 400 output tokens.