Mistral Medium 3.5
Mistral·
Overview
128B dense flagship with 256k context, multimodal (text + vision), open weights under a modified MIT license. First Mistral flagship trained on the new 13,800-GPU Paris facility. Merges the Magistral reasoning line and the Devstral 2 coding line into one model. Mistral did not publish GPQA/HLE/MMLU-Pro at launch; granular third-party scores still pending. Independent composite available: AA Intelligence Index v4.0 = 39 (#2 in class), with very verbose output (~90M tokens vs. 16M class average).
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| SWE-bench Verified | 77.6% | vendor SWE-Bench Verified 77.6% reported by Mistral. |
- GPQA Diamond: not reported by the vendor
- Humanity’s Last Exam (no tools): not reported by the vendor
- SWE-bench Pro: not reported by the vendor
- MMLU: not reported by the vendor
- MMLU-Pro: not reported by the vendor
- Terminal-Bench 2.x: not reported by the vendor
- MMMU-Pro: not reported by the vendor
API pricing
$1.5 per million input tokens, $7.5 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-05-11)
Architecture and hardware
- Parameters
- 128B total, 128B active per token (dense)
- Estimated VRAM at Q4
- ~74 GB, Workstation class
Links
More from Mistral
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.