Wait Which Model?
← Back to directory
NVIDIA

Nemotron 3 Super

Released Mar 11, 2026 · knowledge cutoff unpublished

Status
Superseded
Location
United States
Modality
Text
Context window
1M
Max output
Speed
Price ($/MTok in / out)
$0.10 / $0.50
Cost per task
$0.20
Open weights
Yes
Benchmarks4 of 8 reported

MMLU-Pro

83.73%

SWE-bench Verified

60.47%

AIME (latest)

90.21%

Humanity's Last Exam

18.26%

GPQA Diamond

Terminal-Bench 2.1

LMArena Elo

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Among the fastest open models at its capability — roughly 450 output tokens/sec on hosted endpoints
  • Activates about a tenth of its parameters per token, so it serves far cheaper than its size implies
  • Handles very long inputs without the slowdown dense models hit, thanks to the hybrid state-space design

Weaknesses

  • Text-only, so no screenshots, diagrams or PDFs
  • Reasoning depth well short of frontier flagships — a fast workhorse rather than a hard-problem model

Middle of the Nemotron 3 family; 120B total / 12.7B active MoE. MMLU-Pro 83.73, SWE-bench Verified 60.47, AIME 2025 90.21 and HLE 18.26 per llm-stats' launch analysis. GPQA Diamond left null deliberately: the reported 79.23 is labelled "GPQA no tools", which NVIDIA's own report treats as distinct from GPQA Diamond. Pricing is DeepInfra's rate; OpenRouter lists $0.085/$0.40 and the cross-provider median is about $0.30/$0.80. Artificial Analysis Intelligence Index 36.0.

For developers

API model strings

Not researched

Licence

Not researched

Retirement

No retirement announced

Lineage

No recorded predecessor

News