MiMo-V2.5-Pro
Xiaomi·
LLMsopen-weightFrontier
Overview
Open-weights flagship from Xiaomi. 1.02T total / 42B active hybrid MoE with interleaved Sliding Window + Global Attention (6:1, 128-token window). 1M context, MIT-licensed for commercial use and fine-tuning without additional authorization. Day-0 SGLang/vLLM support. Reports SWE-bench Pro 57.2 and ClawEval 64% Pass³ at ~70K tokens per trajectory.
Capabilities and innovations
1.02T total / 42B active (hybrid MoE)1M context windowHybrid SWA + Global Attention (6:1, 128 window)FP8 E4M3 native precisionMIT license (commercial use + fine-tuning)Interleaved SWA + Global Attention (6:1 ratio)Lightweight MTP modules with dense FFNsDay-0 hardware adaptation (Trainium2, ROCm, Kunlun, T-Head, Enflame, Muxi)Largest fully open-sourced MoE model at release (1.02T total / 42B active, MIT license)
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| SWE-bench Pro | 57.2% | unsourced |
Architecture and hardware
- Parameters
- 1020B total, 42B active per token (MoE)
- Estimated VRAM at Q4
- ~587 GB, Frontier class
- Quantization formats
- FP8
- Recommended runtime
- vLLM
- License
- MIT
Links
More from Xiaomi
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.