AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

MAI-Code-1-Flash

Microsoft·

LLMscodingcloud

Overview

Microsoft's first in-house coding model, part of the MAI family unveiled at Build 2026 and built on Maia 200 silicon without distillation from other labs. Live in GitHub Copilot and VS Code, and available to developers via OpenRouter, Fireworks and Baseten with tunable weights. Vendor-reported benchmarks: SWE-bench Pro 51.2, GPQA Diamond 84.6 (model card).

Capabilities and innovations

Agentic codingGitHub Copilot integrationLow token usageFirst Microsoft in-house coding modelTrained on Maia 200 siliconDeveloper-tunable weights

Benchmarks

BenchmarkScoreSource
GPQA Diamond84.6%vendor
GPQA Diamond 84.6 — MAI-Code-1-Flash model card.
SWE-bench Pro51.2%vendor
SWE-bench Pro 51.2 — Microsoft announcement (vs 35.2 for Claude Haiku 4.5 per their comparison).
  • Humanity’s Last Exam (no tools): not reported by the vendor
  • SWE-bench Verified: not reported by the vendor — Microsoft reports SWE-bench Pro; a Verified figure is referenced for token efficiency but not given as a clean pass rate.
  • MMLU: not reported by the vendor
  • MMLU-Pro: not reported by the vendor
  • Terminal-Bench 2.x: independent evaluation pending — Microsoft cites Terminal-Bench 2 results without a transcribed score.
  • MMMU-Pro: not reported by the vendor

Links

More from Microsoft

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.