Wait Which Model?
← Back to directory
NVIDIA

Nemotron 3.5 Lightning

Released Aug 11, 2026 · knowledge cutoff 2025-09

Status
Superseded
Location
United States
Modality
Text
Context window
1M
Max output
131K
Speed
293 tok/s · 1.12s to first answer token
Price ($/MTok in / out)
$0.08 / $0.20
Cost per task
$0.060
Open weights
Yes
Benchmarks5 of 8 reported

MMLU-Pro

81.62%

GPQA Diamond

75.57%

SWE-bench Verified

52.8%

Terminal-Bench 2.1

23.46%

Humanity's Last Exam

10.47%

AIME (latest)

LMArena Elo

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Answers begin almost immediately and stream fast enough that an agent loop feels interactive rather than batched
  • Interleaved Mamba-2 layers hold throughput up as the context fills, where dense attention models slow down
  • Ships weights, training data and post-training recipes together, so it can genuinely be retrained rather than only run
  • Fits on a single H100 and the licence permits commercial deployment without negotiation

Weaknesses

  • Extremely verbose on reasoning work — it burns roughly twice the tokens of a median model to reach an answer, which eats into the cost advantage its per-token price suggests
  • Text-only: no images, screenshots or documents
  • Built for throughput rather than depth, and falls away sharply on hard reasoning and long-horizon coding

30B total / 3B active MoE with a hybrid Mamba-2 + MoE + attention architecture, released alongside NeMo Switchyard, NVIDIA's open model-routing library. Benchmarks are NVIDIA's own model-card figures for the NVFP4 checkpoint (MMLU-Pro 81.62, GPQA Diamond 75.57, SWE-bench Verified 52.80, Terminal-Bench 2.1 23.46, HLE text-only no-tools 10.47); the full-precision repo returned HTTP 401 so those were not separately confirmed. AIME, LMArena Elo and ARC-AGI-2 are unpublished. Knowledge cutoff is the pre-training date (September 2025); post-training data runs to May 2026. NVIDIA publishes no first-party per-token price — $0.08/$0.20 is OpenRouter's standard route, with DeepInfra at $0.05/$0.20 and CoreWeave at $0.10/$0.25, and a free route also exists. Artificial Analysis Intelligence Index 24 at $0.06 per task; it generated 100M tokens running the index against a 42M median, which is the basis for the verbosity note. OpenRouter describes it as distilled from Nemotron 3 Ultra, but no primary source states it replaces any model, so predecessorId stays null.

For developers

API model strings

  • nvidianvidia/nemotron-3.5-lightning-30b-a3b

Licence

OpenMDW-1.1 · Permissive · commercial use permitted

Retirement

No retirement announced

Lineage

No recorded predecessor

News