Mixtral 8x7B
Mistral·
LLMsopen-weightWorkstation
Overview
First mainstream open Mixture-of-Experts model. Used 8 expert networks with only 2 active per token, achieving 70B-quality at 13B inference cost.
Capabilities and innovations
46.7B total / 12.9B active parameters8 expert networks, 2 active per token32K context windowMultilingual (EN, FR, IT, DE, ES)Sparse Mixture of Experts (SMoE)Expert routing with top-2 gatingEfficient inference via sparse activation
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| MMLU | 70.6% | unsourced |
Architecture and hardware
- Parameters
- 46.7B total, 12.9B active per token (MoE)
- Estimated VRAM at Q4
- ~27 GB, Workstation class
- Quantization formats
- GGUF, GPTQ, AWQ
- Recommended runtime
- Ollama
- License
- Apache 2.0
Links
More from Mistral
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.