AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

K2 Horizon 375B-A23B

IFM·

LLMsopen-weightlocalMulti-GPU Self-Host

Overview

Part of the six-model K2 Horizon fleet released by MBZUAI's Institute of Foundation Models on 2026-09-03 (0.9B, 3.7B, 7B, 32B dense plus MoVA-36B-A4B and 375B-A23B MoE), all Apache 2.0 with weights, intermediate checkpoints, training code, configs, logs and data recipes. Every size is pretrained on roughly 20 trillion tokens, about 10 trillion of them synthetic, with nearly 17 percent explicit problem-solving trajectories. Day-zero support in vLLM, SGLang and Ollama. This entry is the flagship: 375B total capacity, ~23B activated per token, the fleet's strongest model and IFM's enterprise tier. The three sizes without their own entry here are the dense 32B (GPQA Diamond 82.3, Terminal-Bench 2.1 36.6), the 7B (SWE-bench Verified 70.6, BrowseComp 59.0) and the 0.9B (AIME 2026 48.5, HumanEval+ 79.9), built for watches and glasses.

Capabilities and innovations

375B total / 23B active MoE flagshipAgentic tool use, terminal and long-horizon workflowsSix sizes on one shared architecture and tokenizerApache 2.0 licenseMoVA (Mixture-of-Value Attention) — expert routing extended from feed-forward layers into multi-head attentionUno Diffusion — LoRA adapters that decode token blocks in parallel while the autoregressive weights stay frozenIntermediate checkpoints, training logs and data recipes published for every stage including agentic post-trainingSelf-audited its own Terminal-Bench 2.1 launch score for reward hacking and published the downward correction

Benchmarks

BenchmarkScoreSource
GPQA Diamond87.3%vendor
GPQA Diamond 87.3 for K2-Horizon-375B-A23B, launch-post table 'Full Results'. Sibling sizes in their own tables: 32B 82.3, MoVA-36B-A4B 80.8, 3.7B 65.4, 0.9B 27.3 (the 7B table has no GPQA row).
Humanity’s Last Exam (no tools)32%vendor
Humanity's Last Exam without tools, 32.0 for the 375B-A23B. Siblings: MoVA-36B-A4B 25.2, 32B 22.8, 7B 18.6, 3.7B 12.9.
SWE-bench Pro42.6%vendor
SWE-bench Pro 42.6 for the 375B-A23B, run strict (no internet) per the table footnote. Only the flagship table carries this row.
Terminal-Bench 2.x66.9%vendor
Terminal-Bench 2.1. The table prints 70.2; the post's 'From Open Source to Open Science' section reports the audit: 89 tasks x 8 attempts = 712 trials, 500 passing (70.2%), every passing trial then re-checked with Artificial Analysis's harbor analyze tool under the reward_hacking criterion and their verbatim rubric, judged by Codex gpt-5.6-sol. 24 trials across 10 tasks flagged, 79 tasks fully clean, corrected accuracy 66.9 (-3.37 points). The audited figure is recorded here because it is the one the source stands behind. Caveat for cross-reading: the rest of this column is un-audited, so K2 is the only entry measured under the stricter procedure — IFM cites AA flag rates of 2.2% (Claude Fable 5) and 4.1% (GPT-5.6 Luna) for scale. Sibling sizes, un-audited: MoVA-36B-A4B 58.6, 7B 39.1, 32B 36.6, 3.7B 25.1.
  • MMLU-Pro: not reported by the vendor — The launch tables run agentic, coding, scientific-reasoning and general rows (GDPVal-AA, tau3-Banking, Toolathlon, MCPMark, BrowseComp, SciCode, CritPt, AA-LCR, AA-Omniscience). No MMLU or MMLU-Pro value is published for any size.
  • MMLU: not reported by the vendor — Not in any of the six launch tables.
  • SWE-bench Verified: not reported by the vendor — SWE-bench Verified is reported for the 7B (70.6) and 3.7B (68.6) only; the flagship's coding rows are Terminal-Bench 2.1, SciCode, SWE-Atlas-QnA and SWE-bench Pro. Values from a different size are not carried onto this entry.
  • Terminal-Bench 3.0: not reported by the vendor — The launch measured Terminal-Bench 2.1 only.
  • MMMU-Pro: not reported by the vendor — Text-only fleet; no multimodal row in any table.
  • DeepSWE v1.1: not reported by the vendor — No DeepSWE row in any of the six K2 Horizon launch tables; the coding rows are Terminal-Bench 2.1, SciCode, SWE-Atlas-QnA, SWE-bench Pro and SWE-bench Verified.

Architecture and hardware

Parameters
375B total, 23B active per token (MoE)
Estimated VRAM at Q4
~216 GB, Multi-GPU Self-Host class
Quantization formats
GGUF
Recommended runtime
vLLM
License
Apache 2.0

Links

More from IFM

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.