Nemotron 3 Ultra (free) pricing & benchmarks
Nemotron 3 Ultra (free) is available from NVIDIA at Free per million input tokens and Free per million output tokens (Free blended at 3:1). It accepts up to 1,000,000 tokens of context. NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Pricing
Per million tokens ยท OpenRouter
InputFree
OutputFree
Blended (3:1)Free
Specification
nvidia/nemotron-3-ultra-550b-a55b:free
Context window1M
Max output66K
ReleasedJun 2026
LicenceOpen weights
Modalitiesin:text, out:text
No independently-run benchmark scores are published for Nemotron 3 Ultra (free) yet. We show nothing rather than reprinting an unverified figure.
Compare Nemotron 3 Ultra (free)
Nemotron 3 Ultra (free) vs Qwen3.8 Max
Nemotron 3 Ultra (free) vs DeepSeek V4 Flash 0731
Nemotron 3 Ultra (free) vs Claude Opus 5
Nemotron 3 Ultra (free) vs Gemini 3.6 Flash
Nemotron 3 Ultra (free) vs Kimi K3
Nemotron 3 Ultra (free) vs GPT-5.6 Luna
Nemotron 3 Ultra (free) vs GPT-5.6 Terra
Nemotron 3 Ultra (free) vs GPT-5.6 Sol
FAQ
How much does Nemotron 3 Ultra (free) cost?
Nemotron 3 Ultra (free) costs Free per million input tokens and Free per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to Free per million tokens.
What is Nemotron 3 Ultra (free)'s context window?
Nemotron 3 Ultra (free) accepts up to 1,000,000 tokens of context, and can generate up to 65,536 output tokens.