MiMo-V2.5-Pro
Xiaomi·
Overview
Xiaomi's flagship coding/agentic open-weights model. 1.02T total / 42B active hybrid MoE (70 layers, 384 routed experts, 8 active per token) with interleaved Sliding Window + Global Attention at a 6:1 ratio and 128-token window for long-context efficiency. 1M context, MIT licensed, day-0 SGLang/vLLM support. Reports SWE-bench Pro 57.2, GDPVal-AA Elo 1581, ClawEval 64% Pass³ at ~70K tokens per trajectory, and a SysY compiler-in-Rust task solved 233/233 in 4.3h with 672 tool calls. Project lead Fuli Luo (ex-DeepSeek).
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| SWE-bench Pro | 57.2% | third party SWE-bench Pro 57.2 reported via VentureBeat (Build Fast with AI as secondary source). Not corroborated on mimo.xiaomi.com or HuggingFace model card at intake time. |
| MMLU-Pro | 84.59% | unsourced |
- GPQA Diamond: not reported by the vendor
- Humanity’s Last Exam (no tools): not reported by the vendor
- SWE-bench Verified: not reported by the vendor
- MMLU: not reported by the vendor
- Terminal-Bench 2.x: not reported by the vendor
- MMMU-Pro: not reported by the vendor
API pricing
$1 per million input tokens, $3 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-04-29)
Architecture and hardware
- Parameters
- 1020B total, 42B active per token (MoE)
Links
More from Xiaomi
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.