Mistral Medium 3.5
Mistral·
LLMsopen-weightWorkstation
Overview
128B dense flagship from Mistral with 256k context, multimodal (text + vision), open weights under a modified MIT license. Self-hostable on 4 GPUs at ~70GB VRAM with Q4 quantization. First Mistral flagship trained on the new 13,800-GPU Paris facility, merging Magistral (reasoning) and Devstral 2 (coding) into one model.
Capabilities and innovations
128B dense parameters256k context windowMultimodal (text + vision)Reasoning + coding unifiedModified MIT licenseFirst flagship trained on Mistral's 13,800-GPU Paris facilityUnified Magistral (reasoning) + Devstral 2 (coding) in one modelModified MIT release at flagship scale
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| SWE-bench Verified | 77.6% | unsourced |
Architecture and hardware
- Parameters
- 128B total, 128B active per token (Dense)
- Estimated VRAM at Q4
- ~74 GB, Workstation class
- Quantization formats
- GGUF, GPTQ, AWQ
- Recommended runtime
- vLLM
- License
- Modified MIT
Links
More from Mistral
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.