Llama 4 Scout
Meta·
LLMsopen-weightWorkstation
Overview
Open MoE multimodal model with 16 experts. 10M token context window via iRoPE. Efficient inference at 17B active parameters.
Capabilities and innovations
109B total / 17B active parameters (MoE)16 experts, 1 active per token10M token context windowNative multimodal (text + image)Interleaved early-fusion for multimodalityMetaP for hyperparameter predictionMixture of Experts at massive scaleiRoPE architecture for long context
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 57.2% | unsourced |
| Humanity’s Last Exam (no tools) | 14.7% | unsourced |
| MMLU | 85.4% | unsourced |
Architecture and hardware
- Parameters
- 109B total, 17B active per token (MoE)
- Estimated VRAM at Q4
- ~63 GB, Workstation class
- Quantization formats
- GGUF, GPTQ, AWQ, FP8
- Recommended runtime
- vLLM
- License
- Llama 4 Community License
Links
More from Meta
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.