Llama 4 Maverick
Meta·
LLMsopen-weightMulti-GPU Self-Host
Overview
Open MoE multimodal model with 128 experts. Frontier-level quality at efficient inference cost with 17B active parameters.
Capabilities and innovations
400B total / 17B active parameters (MoE)128 experts, 1 active per tokenFrontier reasoning qualityNative multimodal (text + image)Interleaved early-fusion for multimodalityMetaP for hyperparameter predictionMixture of Experts at massive scaleiRoPE architecture for long context
Benchmarks
No benchmark values are recorded for this model.
Architecture and hardware
- Parameters
- 400B total, 17B active per token (MoE)
- Estimated VRAM at Q4
- ~230 GB, Multi-GPU Self-Host class
- Quantization formats
- GGUF, GPTQ, AWQ, FP8
- Recommended runtime
- vLLM
- License
- Llama 4 Community License
Links
More from Meta
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.