Nemotron 3.5 Lightning
Released Aug 11, 2026 · knowledge cutoff 2025-09
- Status
- Superseded
- Location
- United States
- Modality
- Text
- Context window
- 1M
- Max output
- 131K
- Speed
- 293 tok/s · 1.12s to first answer token
- Price ($/MTok in / out)
- $0.08 / $0.20
- Cost per task
- $0.060
- Open weights
- Yes
Benchmarks5 of 8 reported
MMLU-Pro
GPQA Diamond
SWE-bench Verified
Terminal-Bench 2.1
Humanity's Last Exam
AIME (latest)
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- Answers begin almost immediately and stream fast enough that an agent loop feels interactive rather than batched
- Interleaved Mamba-2 layers hold throughput up as the context fills, where dense attention models slow down
- Ships weights, training data and post-training recipes together, so it can genuinely be retrained rather than only run
- Fits on a single H100 and the licence permits commercial deployment without negotiation
Weaknesses
- Extremely verbose on reasoning work — it burns roughly twice the tokens of a median model to reach an answer, which eats into the cost advantage its per-token price suggests
- Text-only: no images, screenshots or documents
- Built for throughput rather than depth, and falls away sharply on hard reasoning and long-horizon coding
30B total / 3B active MoE with a hybrid Mamba-2 + MoE + attention architecture, released alongside NeMo Switchyard, NVIDIA's open model-routing library. Benchmarks are NVIDIA's own model-card figures for the NVFP4 checkpoint (MMLU-Pro 81.62, GPQA Diamond 75.57, SWE-bench Verified 52.80, Terminal-Bench 2.1 23.46, HLE text-only no-tools 10.47); the full-precision repo returned HTTP 401 so those were not separately confirmed. AIME, LMArena Elo and ARC-AGI-2 are unpublished. Knowledge cutoff is the pre-training date (September 2025); post-training data runs to May 2026. NVIDIA publishes no first-party per-token price — $0.08/$0.20 is OpenRouter's standard route, with DeepInfra at $0.05/$0.20 and CoreWeave at $0.10/$0.25, and a free route also exists. Artificial Analysis Intelligence Index 24 at $0.06 per task; it generated 100M tokens running the index against a 42M median, which is the basis for the verbosity note. OpenRouter describes it as distilled from Nemotron 3 Ultra, but no primary source states it replaces any model, so predecessorId stays null.
For developers
API model strings
- nvidia
nvidia/nemotron-3.5-lightning-30b-a3b
Licence
OpenMDW-1.1 · Permissive · commercial use permitted
Retirement
No retirement announced
Lineage
No recorded predecessor