AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Llama 3.1 405B

Meta·

LLMsopen-weightMulti-GPU Self-Host

Overview

First open-weight frontier-class model. The 405B variant matched GPT-4 on many benchmarks, proving open models could compete at the highest tier.

Capabilities and innovations

8B / 70B / 405B parameters128K context windowTool use and function callingMultilingual (8 languages)FP8 quantization for training128K context via progressive trainingIterative DPO alignment

Benchmarks

BenchmarkScoreSource
GPQA Diamond50.7%unsourced
SWE-bench Verified33.2%unsourced
MMLU88.6%unsourced

Architecture and hardware

Parameters
405B total (Dense)
Estimated VRAM at Q4
~233 GB, Multi-GPU Self-Host class
Quantization formats
GGUF, GPTQ, AWQ, FP8
Recommended runtime
Ollama
License
Llama 3.1 Community License

Links

More from Meta

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.