Kimi K3
Moonshot AI·
LLMsopen-weightcloud
Overview
2.8T MoE (~50B active, 16 of 896 experts) native multimodal flagship with a 1M-token context window. Launched via API, app and playground on 2026-07-16; open weights scheduled for 2026-07-27.
Capabilities and innovations
2.8T total / ~50B active MoE (896 experts, 16 routed)1M token context windowNatively multimodal (text, image, video)Vision-in-the-loop (screenshot inspect + code edit)Open weights scheduled 2026-07-27Kimi Delta Attention2.8T Open-Weight MoE (largest open-weight at launch)Vision-in-the-Loop Agent Feedback1M-Token Native Context
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 93.5% | vendor GPQA-Diamond 93.5 (official launch blog, 2026-07-16). |
| Humanity’s Last Exam (no tools) | 43.5% | vendor HLE-Full 43.5 no-tools (text-only), recorded for cross-model comparability. Moonshot also reports 56.0 with tools, which is not comparable to the no-tools HLE used here. |
| MMLU-Pro | 87.97% | independent MMLU Pro 87.97%, measured by Vals AI (accessed 2026-08-17). Moonshot publishes no MMLU-Pro value. Vals AI measures GPQA Diamond at 92.93 on the same run, 0.6 points below Moonshot's self-reported 93.5, so this harness reads in line with the vendor table. |
| Terminal-Bench 2.x | 88.3% | vendor Terminal-Bench 2.1 88.3 (official launch blog). |
| MMMU-Pro | 81.6% | vendor MMMU-Pro 81.6 (official launch blog). |
- SWE-bench Verified: not reported by the vendor
- SWE-bench Pro: not reported by the vendor — Moonshot reports its own coding suite (Program Bench, SWE Marathon, FrontierSWE, DeepSWE) instead of SWE-bench Pro.
API pricing
$3 per million input tokens, $15 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-07-17)
Architecture and hardware
- Parameters
- 2800B total, 50B active per token (MoE)
Links
More from Moonshot AI
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.