AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

MiMo-V2.6-Flash

Xiaomi·

LLMsopen-weightMulti-GPU Self-Host

Overview

Smaller sibling of MiMo-V2.6-Pro, released the same day: 309B total / 15B active MoE (48 layers, 256 experts, top-8 routing) with hybrid sliding-window and global attention. 1M-token context, MIT license. API $0.14 / $0.28 per million input/output tokens.

Capabilities and innovations

309B total / 15B active (MoE)1M token context windowHybrid SWA + global attentionMIT licenseCommunity GGUF, MLX and NVFP4 quantizations at launch

Benchmarks

BenchmarkScoreSource
Terminal-Bench 2.x76.4%independent
Terminal-Bench 2.1 76.40%, measured by Vals AI (accessed 2026-09-23), above the larger Pro model on the same harness. Xiaomi self-reports 87.6.
  • GPQA Diamond: independent evaluation pending
  • Humanity’s Last Exam (no tools): independent evaluation pending — Not yet on Artificial Analysis (checked 2026-09-23).
  • SWE-bench Pro: not reported by the vendor
  • MMLU-Pro: independent evaluation pending
  • Terminal-Bench 3.0: not reported by the vendor — Not on the official tbench.ai Terminal-Bench 3.0 leaderboard (checked 2026-09-23); 4.0 is the current release.

Architecture and hardware

Parameters
309B total, 15B active per token (MoE)
Estimated VRAM at Q4
~178 GB, Multi-GPU Self-Host class
Quantization formats
BF16, FP8, GGUF, MLX, NVFP4
Recommended runtime
SGLang
License
MIT

Links

More from Xiaomi

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.