AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Nemotron 3 Ultra

NVIDIA·

LLMsopen-weightFrontier

Overview

Largest model of the Nemotron 3 family: 550B total / 55B active hybrid Mamba-Transformer MoE. Released with weights, training data and recipes under OpenMDW-1.1.

Capabilities and innovations

550B total / 55B active (MoE)Hybrid Mamba-Transformer262K context (1M with NVFP4)OpenMDW-1.1 licenseHybrid Mamba-Transformer MoEOpen training data and recipes (OpenMDW-1.1)

Benchmarks

No benchmark values are recorded for this model.

  • GPQA Diamond: independent evaluation pending
  • Humanity’s Last Exam (no tools): independent evaluation pending
  • SWE-bench Pro: independent evaluation pending
  • Terminal-Bench 2.x: independent evaluation pending
  • MMLU-Pro: independent evaluation pending
  • MMMU-Pro: independent evaluation pending

Architecture and hardware

Parameters
550B total, 55B active per token (MoE)
Estimated VRAM at Q4
~317 GB, Frontier class
Quantization formats
NVFP4, BF16, GGUF
Recommended runtime
vLLM
License
OpenMDW-1.1

Links

More from NVIDIA

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.