AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Mixtral 8x7B

Mistral·

LLMsopen-weightWorkstation

Overview

First mainstream open Mixture-of-Experts model. Used 8 expert networks with only 2 active per token, achieving 70B-quality at 13B inference cost.

Capabilities and innovations

46.7B total / 12.9B active parameters8 expert networks, 2 active per token32K context windowMultilingual (EN, FR, IT, DE, ES)Sparse Mixture of Experts (SMoE)Expert routing with top-2 gatingEfficient inference via sparse activation

Benchmarks

BenchmarkScoreSource
MMLU70.6%unsourced

Architecture and hardware

Parameters
46.7B total, 12.9B active per token (MoE)
Estimated VRAM at Q4
~27 GB, Workstation class
Quantization formats
GGUF, GPTQ, AWQ
Recommended runtime
Ollama
License
Apache 2.0

Links

More from Mistral

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.