AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Mistral Large 4

Mistral·

LLMsopen-weightcloud

Overview

1.05T-parameter MoE (~490B active, plus a 1.6B vision encoder), text and image input, text output, nicknamed "Le Chonk". Announced at AI Everything Abu Dhabi on 2026-10-06 with an API preview; open weights scheduled for 2026-10-27. Mistral states a context of up to 1M tokens, OpenRouter serves 512k. Mistral published no GPQA, HLE, MMLU-Pro or SWE-bench Pro at launch; its table leads with DeepSWE v1.1 plus domain evals (FinWorkBench, Harvey Legal Agent Benchmark, visual grounding).

Capabilities and innovations

1.05T total / ~490B active MoEText + image inputUp to 1M token context (per Mistral)Agentic codingOpen weights scheduled 2026-10-27Largest Mistral model to date

Benchmarks

BenchmarkScoreSource
Terminal-Bench 4.022.73%independent
Terminal-Bench 4.0 22.73% (±0.88), measured by Vals AI (accessed 2026-10-06). Mistral reports 28.3; the gap exceeds 3 points, so the independent value is primary.
DeepSWE v1.161.7%vendor
DeepSWE v1.1 61.7 from Mistral's launch table (headlines round it to 62). Mistral's own model, explicitly v1.1. The same table lists GLM-5.3 at 61, DeepSeek V4 Pro 57, Qwen 3.8 Max 51 and Reflection Beam 44; those are Mistral's measurements of other vendors' models and are not taken over. Read via secondary coverage, mistral.ai was not reachable from the intake sandbox.
  • GPQA Diamond: not reported by the vendor
  • Humanity’s Last Exam (no tools): not reported by the vendor
  • SWE-bench Pro: not reported by the vendor
  • MMLU-Pro: not reported by the vendor
  • Terminal-Bench 2.x: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

API pricing

$0.68 per million input tokens, $2.09 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-10-06)

Architecture and hardware

Parameters
1050B total, 490B active per token (MoE)

Links

More from Mistral

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.