AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Llama 4 Scout

Meta·

LLMsopen-weightWorkstation

Overview

Open MoE multimodal model with 16 experts. 10M token context window via iRoPE. Efficient inference at 17B active parameters.

Capabilities and innovations

109B total / 17B active parameters (MoE)16 experts, 1 active per token10M token context windowNative multimodal (text + image)Interleaved early-fusion for multimodalityMetaP for hyperparameter predictionMixture of Experts at massive scaleiRoPE architecture for long context

Benchmarks

BenchmarkScoreSource
GPQA Diamond57.2%unsourced
Humanity’s Last Exam (no tools)14.7%unsourced
MMLU85.4%unsourced

Architecture and hardware

Parameters
109B total, 17B active per token (MoE)
Estimated VRAM at Q4
~63 GB, Workstation class
Quantization formats
GGUF, GPTQ, AWQ, FP8
Recommended runtime
vLLM
License
Llama 4 Community License

Links

More from Meta

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.