Nemotron 3.5 Lightning
NVIDIA·
LLMsopen-weightEdge / Consumer
Overview
NVIDIA's current open-weight release: a 30B / 3B-active hybrid model interleaving Mamba-2, MoE and select attention layers, tuned for throughput in agent loops rather than peak scores. Text-only with a 1M-token context, shipped under OpenMDW-1.1 together with training data and post-training recipes, in BF16, NVFP4 and community GGUF builds. A 4-bit quant runs on a single 24 GB consumer GPU. NVIDIA reports the largest gains over Nemotron 3 Nano on agentic evaluations; Artificial Analysis measures roughly 670 output tokens/s on pre-release infrastructure.
Capabilities and innovations
30B total / 3B active parameters (MoE)Hybrid Mamba-2 + MoE + attention layers1M token context windowText-only reasoningBF16, NVFP4 and GGUF buildsOpenMDW-1.1 with open training data and recipesMamba-2 / MoE hybrid at 3B active parametersNVFP4 checkpoint with near-BF16 benchmark parityOpen training data, RL environments and post-training recipes
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 75.44% | vendor GPQA Diamond 75.44 from NVIDIA's BF16 model card (the NVFP4 build measures 75.57). Card not fetchable from this environment; value taken from consistent secondary transcriptions of the card. |
| SWE-bench Verified | 51.56% | vendor SWE-bench Verified 51.56 BF16 (NVFP4: 52.80). Era-1 benchmark, recorded only because NVIDIA reports it. |
| MMLU-Pro | 81.94% | vendor MMLU-Pro 81.94 from the same BF16 model card. |
| Terminal-Bench 2.x | 24% | third party Terminal-Bench v2.1 24% measured by Artificial Analysis at launch (Nemotron 3 Nano: 7%). |
- Humanity’s Last Exam (no tools): independent evaluation pending — HLE is part of the Artificial Analysis Intelligence Index run for this model (composite score 24), but no per-benchmark HLE value has been published.
- SWE-bench Pro: not reported by the vendor — NVIDIA's model card reports SWE-bench Verified and PinchBench instead.
- MMMU-Pro: not reported by the vendor — Not applicable — the model is text-only.
Architecture and hardware
- Parameters
- 30B total, 3B active per token (MoE)
- Estimated VRAM at Q4
- ~18 GB, Edge / Consumer class
- Quantization formats
- NVFP4, BF16, GGUF
- Recommended runtime
- Ollama
- License
- OpenMDW-1.1
Links
More from NVIDIA
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.