AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Mistral Medium 3.5

Mistral·

LLMsflagshipcloud + localWorkstation

Overview

128B dense flagship with 256k context, multimodal (text + vision), open weights under a modified MIT license. First Mistral flagship trained on the new 13,800-GPU Paris facility. Merges the Magistral reasoning line and the Devstral 2 coding line into one model. Mistral did not publish GPQA/HLE/MMLU-Pro at launch; granular third-party scores still pending. Independent composite available: AA Intelligence Index v4.0 = 39 (#2 in class), with very verbose output (~90M tokens vs. 16M class average).

Capabilities and innovations

128B dense parameters256k context windowMultimodal (text + vision)Reasoning + coding unifiedOpen weights (Modified MIT)Self-hostable on 4 GPUs (~70GB VRAM Q4)First flagship trained on Mistral's 13,800-GPU Paris facilityUnified Magistral (reasoning) + Devstral 2 (coding) in one modelModified MIT release at flagship scale

Benchmarks

BenchmarkScoreSource
SWE-bench Verified77.6%vendor
SWE-Bench Verified 77.6% reported by Mistral.
  • GPQA Diamond: not reported by the vendor
  • Humanity’s Last Exam (no tools): not reported by the vendor
  • SWE-bench Pro: not reported by the vendor
  • MMLU: not reported by the vendor
  • MMLU-Pro: not reported by the vendor
  • Terminal-Bench 2.x: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

API pricing

$1.5 per million input tokens, $7.5 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-05-11)

Architecture and hardware

Parameters
128B total, 128B active per token (dense)
Estimated VRAM at Q4
~74 GB, Workstation class

Links

More from Mistral

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.