MiMo-V2.6-Flash
Xiaomi·
LLMsopen-weightMulti-GPU Self-Host
Overview
Smaller sibling of MiMo-V2.6-Pro, released the same day: 309B total / 15B active MoE (48 layers, 256 experts, top-8 routing) with hybrid sliding-window and global attention. 1M-token context, MIT license. API $0.14 / $0.28 per million input/output tokens.
Capabilities and innovations
309B total / 15B active (MoE)1M token context windowHybrid SWA + global attentionMIT licenseCommunity GGUF, MLX and NVFP4 quantizations at launch
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| Terminal-Bench 2.x | 76.4% | independent Terminal-Bench 2.1 76.40%, measured by Vals AI (accessed 2026-09-23), above the larger Pro model on the same harness. Xiaomi self-reports 87.6. |
- GPQA Diamond: independent evaluation pending
- Humanity’s Last Exam (no tools): independent evaluation pending — Not yet on Artificial Analysis (checked 2026-09-23).
- SWE-bench Pro: not reported by the vendor
- MMLU-Pro: independent evaluation pending
- Terminal-Bench 3.0: not reported by the vendor — Not on the official tbench.ai Terminal-Bench 3.0 leaderboard (checked 2026-09-23); 4.0 is the current release.
Architecture and hardware
- Parameters
- 309B total, 15B active per token (MoE)
- Estimated VRAM at Q4
- ~178 GB, Multi-GPU Self-Host class
- Quantization formats
- BF16, FP8, GGUF, MLX, NVFP4
- Recommended runtime
- SGLang
- License
- MIT
Links
More from Xiaomi
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.