MAI-Code-1-Flash
Microsoft·
LLMscodingcloud
Overview
Microsoft's first in-house coding model, part of the MAI family unveiled at Build 2026 and built on Maia 200 silicon without distillation from other labs. Live in GitHub Copilot and VS Code, and available to developers via OpenRouter, Fireworks and Baseten with tunable weights. Vendor-reported benchmarks: SWE-bench Pro 51.2, GPQA Diamond 84.6 (model card).
Capabilities and innovations
Agentic codingGitHub Copilot integrationLow token usageFirst Microsoft in-house coding modelTrained on Maia 200 siliconDeveloper-tunable weights
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 84.6% | vendor GPQA Diamond 84.6 — MAI-Code-1-Flash model card. |
| SWE-bench Pro | 51.2% | vendor SWE-bench Pro 51.2 — Microsoft announcement (vs 35.2 for Claude Haiku 4.5 per their comparison). |
- Humanity’s Last Exam (no tools): not reported by the vendor
- SWE-bench Verified: not reported by the vendor — Microsoft reports SWE-bench Pro; a Verified figure is referenced for token efficiency but not given as a clean pass rate.
- MMLU: not reported by the vendor
- MMLU-Pro: not reported by the vendor
- Terminal-Bench 2.x: independent evaluation pending — Microsoft cites Terminal-Bench 2 results without a transcribed score.
- MMMU-Pro: not reported by the vendor
Links
More from Microsoft
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.