Nemotron 3 Ultra
NVIDIA·
LLMsopen-weightFrontier
Overview
Largest model of the Nemotron 3 family: 550B total / 55B active hybrid Mamba-Transformer MoE. Released with weights, training data and recipes under OpenMDW-1.1.
Capabilities and innovations
550B total / 55B active (MoE)Hybrid Mamba-Transformer262K context (1M with NVFP4)OpenMDW-1.1 licenseHybrid Mamba-Transformer MoEOpen training data and recipes (OpenMDW-1.1)
Benchmarks
No benchmark values are recorded for this model.
- GPQA Diamond: independent evaluation pending
- Humanity’s Last Exam (no tools): independent evaluation pending
- SWE-bench Pro: independent evaluation pending
- Terminal-Bench 2.x: independent evaluation pending
- MMLU-Pro: independent evaluation pending
- MMMU-Pro: independent evaluation pending
Architecture and hardware
- Parameters
- 550B total, 55B active per token (MoE)
- Estimated VRAM at Q4
- ~317 GB, Frontier class
- Quantization formats
- NVFP4, BF16, GGUF
- Recommended runtime
- vLLM
- License
- OpenMDW-1.1
Links
More from NVIDIA
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.