AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Mistral 7B

Mistral·

LLMsopen-weightEdge / Consumer

Overview

Efficiency breakthrough from Mistral AI. A 7B model that outperformed Llama 2 13B on all benchmarks using Grouped-Query Attention and Sliding Window Attention.

Capabilities and innovations

7B parametersOutperforms Llama 2 13B8K context with sliding windowApache 2.0 licenseSliding Window Attention (SWA)Grouped-Query Attention (GQA)Rolling buffer KV cache

Benchmarks

BenchmarkScoreSource
MMLU60.1%unsourced

Architecture and hardware

Parameters
7B total (Dense)
Estimated VRAM at Q4
~5 GB, Edge / Consumer class
Quantization formats
GGUF, GPTQ, AWQ
Recommended runtime
Ollama
License
Apache 2.0

Links

More from Mistral

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.