AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Kimi K3

Moonshot AI·

LLMsopen-weightcloud

Overview

2.8T MoE (~50B active, 16 of 896 experts) native multimodal flagship with a 1M-token context window. Launched via API, app and playground on 2026-07-16; open weights scheduled for 2026-07-27.

Capabilities and innovations

2.8T total / ~50B active MoE (896 experts, 16 routed)1M token context windowNatively multimodal (text, image, video)Vision-in-the-loop (screenshot inspect + code edit)Open weights scheduled 2026-07-27Kimi Delta Attention2.8T Open-Weight MoE (largest open-weight at launch)Vision-in-the-Loop Agent Feedback1M-Token Native Context

Benchmarks

BenchmarkScoreSource
GPQA Diamond93.5%vendor
GPQA-Diamond 93.5 (official launch blog, 2026-07-16).
Humanity’s Last Exam (no tools)43.5%vendor
HLE-Full 43.5 no-tools (text-only), recorded for cross-model comparability. Moonshot also reports 56.0 with tools, which is not comparable to the no-tools HLE used here.
MMLU-Pro87.97%independent
MMLU Pro 87.97%, measured by Vals AI (accessed 2026-08-17). Moonshot publishes no MMLU-Pro value. Vals AI measures GPQA Diamond at 92.93 on the same run, 0.6 points below Moonshot's self-reported 93.5, so this harness reads in line with the vendor table.
Terminal-Bench 2.x88.3%vendor
Terminal-Bench 2.1 88.3 (official launch blog).
MMMU-Pro81.6%vendor
MMMU-Pro 81.6 (official launch blog).
  • SWE-bench Verified: not reported by the vendor
  • SWE-bench Pro: not reported by the vendor — Moonshot reports its own coding suite (Program Bench, SWE Marathon, FrontierSWE, DeepSWE) instead of SWE-bench Pro.

API pricing

$3 per million input tokens, $15 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-07-17)

Architecture and hardware

Parameters
2800B total, 50B active per token (MoE)

Links

    More from Moonshot AI

    Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.