AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Nemotron 3.5 Lightning

NVIDIA·

LLMsopen-weightEdge / Consumer

Overview

NVIDIA's current open-weight release: a 30B / 3B-active hybrid model interleaving Mamba-2, MoE and select attention layers, tuned for throughput in agent loops rather than peak scores. Text-only with a 1M-token context, shipped under OpenMDW-1.1 together with training data and post-training recipes, in BF16, NVFP4 and community GGUF builds. A 4-bit quant runs on a single 24 GB consumer GPU. NVIDIA reports the largest gains over Nemotron 3 Nano on agentic evaluations; Artificial Analysis measures roughly 670 output tokens/s on pre-release infrastructure.

Capabilities and innovations

30B total / 3B active parameters (MoE)Hybrid Mamba-2 + MoE + attention layers1M token context windowText-only reasoningBF16, NVFP4 and GGUF buildsOpenMDW-1.1 with open training data and recipesMamba-2 / MoE hybrid at 3B active parametersNVFP4 checkpoint with near-BF16 benchmark parityOpen training data, RL environments and post-training recipes

Benchmarks

BenchmarkScoreSource
GPQA Diamond75.44%vendor
GPQA Diamond 75.44 from NVIDIA's BF16 model card (the NVFP4 build measures 75.57). Card not fetchable from this environment; value taken from consistent secondary transcriptions of the card.
SWE-bench Verified51.56%vendor
SWE-bench Verified 51.56 BF16 (NVFP4: 52.80). Era-1 benchmark, recorded only because NVIDIA reports it.
MMLU-Pro81.94%vendor
MMLU-Pro 81.94 from the same BF16 model card.
Terminal-Bench 2.x24%third party
Terminal-Bench v2.1 24% measured by Artificial Analysis at launch (Nemotron 3 Nano: 7%).
  • Humanity’s Last Exam (no tools): independent evaluation pending — HLE is part of the Artificial Analysis Intelligence Index run for this model (composite score 24), but no per-benchmark HLE value has been published.
  • SWE-bench Pro: not reported by the vendor — NVIDIA's model card reports SWE-bench Verified and PinchBench instead.
  • MMMU-Pro: not reported by the vendor — Not applicable — the model is text-only.

Architecture and hardware

Parameters
30B total, 3B active per token (MoE)
Estimated VRAM at Q4
~18 GB, Edge / Consumer class
Quantization formats
NVFP4, BF16, GGUF
Recommended runtime
Ollama
License
OpenMDW-1.1

Links

More from NVIDIA

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.