MAI-Thinking-1
Microsoft·
LLMsreasoningcloud
Overview
Microsoft AI's first flagship reasoning model, a medium-sized model in the first-party MAI family launched at Build 2026. Distributed via OpenRouter, Fireworks and Baseten.
Capabilities and innovations
Extended reasoningAgentic codingMathematical reasoningTool useFirst-party Microsoft reasoning flagshipMAI model family (Build 2026)
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 84.2% | third party GPQA-Diamond 84.2 (aggregator; corroborated by llmreference). Not reported with a number in Microsoft's launch blog; independent vals.ai evaluation pending. |
| SWE-bench Pro | 52.8% | third party SWE-bench Pro 52.8 (aggregator); Microsoft reported ~53% at Build 2026, on par with Claude Opus 4.6. Not yet on vals.ai for independent confirmation. |
- Humanity’s Last Exam (no tools): independent evaluation pending
- MMLU-Pro: independent evaluation pending
- Terminal-Bench 2.x: independent evaluation pending
- MMMU-Pro: independent evaluation pending
Links
More from Microsoft
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.