AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Llama 4 Maverick

Meta·

LLMsopen-weightMulti-GPU Self-Host

Overview

Open MoE multimodal model with 128 experts. Frontier-level quality at efficient inference cost with 17B active parameters.

Capabilities and innovations

400B total / 17B active parameters (MoE)128 experts, 1 active per tokenFrontier reasoning qualityNative multimodal (text + image)Interleaved early-fusion for multimodalityMetaP for hyperparameter predictionMixture of Experts at massive scaleiRoPE architecture for long context

Benchmarks

No benchmark values are recorded for this model.

Architecture and hardware

Parameters
400B total, 17B active per token (MoE)
Estimated VRAM at Q4
~230 GB, Multi-GPU Self-Host class
Quantization formats
GGUF, GPTQ, AWQ, FP8
Recommended runtime
vLLM
License
Llama 4 Community License

Links

More from Meta

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.