Nemotron 3 Ultra
Released Jun 4, 2026 · knowledge cutoff unpublished
- Status
- Unknown
- Location
- United States
- Modality
- Text
- Context window
- 1M
- Max output
- —
- Speed
- —
- Price ($/MTok in / out)
- $0.50 / $2.20
- Cost per task
- $0.25
- Open weights
- Yes
Benchmarks0 of 8 reported
MMLU-Pro
GPQA Diamond
SWE-bench Verified
Terminal-Bench 2.1
AIME (latest)
Humanity's Last Exam
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- Hybrid Mamba-Transformer routing holds throughput up as context grows, where dense attention models slow down
- Serves faster than most open models near its size — over 400 tokens/sec on hosted endpoints
- Ships training data and recipes alongside the weights, so it can be genuinely retrained rather than only run
- Licence permits commercial deployment and self-hosting without negotiation
Weaknesses
- Text-only — no images, screenshots or documents
- Trails the leading Chinese open-weights models on composite intelligence despite its size
- NVIDIA published no standard academic benchmark table at launch, so it is hard to place precisely
Largest of the three-model Nemotron 3 family (Nano, Super, Ultra); 550B total / 55B active MoE. Artificial Analysis Intelligence Index 47.7 — ahead of Gemma 4 31B (39.2) and Nemotron 3 Super (36.0), behind Kimi K2.6 (53.9). None of the site's tracked benchmarks could be attributed with confidence: NVIDIA's technical report contains many ablation and quantisation tables rather than a single headline model table, so they stay null rather than risk quoting an intermediate checkpoint. Pricing is OpenRouter's listing; NVIDIA publishes no first-party per-token price. Context is 1M per NVIDIA, though OpenRouter serves a 512K effective window.
For developers
API model strings
Not researched
Licence
Not researched
Retirement
No retirement announced
Lineage
No recorded predecessor