Llama 3.1 405B
Meta·
LLMsopen-weightMulti-GPU Self-Host
Overview
First open-weight frontier-class model. The 405B variant matched GPT-4 on many benchmarks, proving open models could compete at the highest tier.
Capabilities and innovations
8B / 70B / 405B parameters128K context windowTool use and function callingMultilingual (8 languages)FP8 quantization for training128K context via progressive trainingIterative DPO alignment
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 50.7% | unsourced |
| SWE-bench Verified | 33.2% | unsourced |
| MMLU | 88.6% | unsourced |
Architecture and hardware
- Parameters
- 405B total (Dense)
- Estimated VRAM at Q4
- ~233 GB, Multi-GPU Self-Host class
- Quantization formats
- GGUF, GPTQ, AWQ, FP8
- Recommended runtime
- Ollama
- License
- Llama 3.1 Community License
Links
More from Meta
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.