Wait Which Model?
← Back to directory
NVIDIA

Nemotron 3 Ultra

Released Jun 4, 2026 · knowledge cutoff unpublished

Status
Unknown
Location
United States
Modality
Text
Context window
1M
Max output
Speed
Price ($/MTok in / out)
$0.50 / $2.20
Cost per task
$0.25
Open weights
Yes
Benchmarks0 of 8 reported

MMLU-Pro

GPQA Diamond

SWE-bench Verified

Terminal-Bench 2.1

AIME (latest)

Humanity's Last Exam

LMArena Elo

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Hybrid Mamba-Transformer routing holds throughput up as context grows, where dense attention models slow down
  • Serves faster than most open models near its size — over 400 tokens/sec on hosted endpoints
  • Ships training data and recipes alongside the weights, so it can be genuinely retrained rather than only run
  • Licence permits commercial deployment and self-hosting without negotiation

Weaknesses

  • Text-only — no images, screenshots or documents
  • Trails the leading Chinese open-weights models on composite intelligence despite its size
  • NVIDIA published no standard academic benchmark table at launch, so it is hard to place precisely

Largest of the three-model Nemotron 3 family (Nano, Super, Ultra); 550B total / 55B active MoE. Artificial Analysis Intelligence Index 47.7 — ahead of Gemma 4 31B (39.2) and Nemotron 3 Super (36.0), behind Kimi K2.6 (53.9). None of the site's tracked benchmarks could be attributed with confidence: NVIDIA's technical report contains many ablation and quantisation tables rather than a single headline model table, so they stay null rather than risk quoting an intermediate checkpoint. Pricing is OpenRouter's listing; NVIDIA publishes no first-party per-token price. Context is 1M per NVIDIA, though OpenRouter serves a 512K effective window.

For developers

API model strings

Not researched

Licence

Not researched

Retirement

No retirement announced

Lineage

No recorded predecessor

News