AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

K2 Horizon MoVA-36B-A4B

IFM·

LLMsopen-weightlocalEdge / Consumer

Overview

Part of the six-model K2 Horizon fleet released by MBZUAI's Institute of Foundation Models on 2026-09-03 (0.9B, 3.7B, 7B, 32B dense plus MoVA-36B-A4B and 375B-A23B MoE), all Apache 2.0 with weights, intermediate checkpoints, training code, configs, logs and data recipes. Every size is pretrained on roughly 20 trillion tokens, about 10 trillion of them synthetic, with nearly 17 percent explicit problem-solving trajectories. Day-zero support in vLLM, SGLang and Ollama. This entry is the MoVA model: 36B total capacity activating only ~4B parameters per token, and the first release of Mixture-of-Value Attention — expert routing extended from the feed-forward layers into multi-head attention itself, compatible with FlashAttention, grouped-query and sparse attention. IFM trained it under the same conditions as the dense 32B as a controlled dense-versus-sparse comparison; it lands slightly below the 32B on scientific reasoning but well ahead of it on agentic terminal use (58.6 against 36.6 on Terminal-Bench 2.1).

Capabilities and innovations

36B total / 4B active MoERuns on a single 24 GB consumer GPU at Q4Agentic tool use and terminal workflowsApache 2.0 licenseMoVA (Mixture-of-Value Attention) — first model to route attention values through expertsTrained under identical conditions to the dense 32B as a controlled architecture comparison

Benchmarks

BenchmarkScoreSource
GPQA Diamond80.8%vendor
GPQA Diamond 80.8, launch-post table 'IFM/K2-Horizon-MoVA-36B-A4B'.
Humanity’s Last Exam (no tools)25.2%vendor
Humanity's Last Exam without tools, 25.2, same table.
Terminal-Bench 2.x58.6%vendor
Terminal-Bench 2.1 58.6, same table. Un-audited: IFM ran its reward-hacking audit on the 375B-A23B only, so this value sits on the same footing as the rest of this column.
  • SWE-bench Pro: not reported by the vendor — The MoVA table carries no SWE-bench row of any generation. Only the 375B-A23B table reports SWE-bench Pro.
  • SWE-bench Verified: not reported by the vendor — SWE-bench Verified is reported for the 7B and 3.7B only.
  • MMLU-Pro: not reported by the vendor — No MMLU or MMLU-Pro row in any of the six launch tables.
  • MMLU: not reported by the vendor — Not in any of the six launch tables.
  • Terminal-Bench 3.0: not reported by the vendor — The launch measured Terminal-Bench 2.1 only.
  • MMMU-Pro: not reported by the vendor — Text-only fleet.
  • DeepSWE v1.1: not reported by the vendor — No DeepSWE row in any of the six K2 Horizon launch tables; the coding rows are Terminal-Bench 2.1, SciCode, SWE-Atlas-QnA, SWE-bench Pro and SWE-bench Verified.

Architecture and hardware

Parameters
36B total, 4B active per token (MoE)
Estimated VRAM at Q4
~21 GB, Edge / Consumer class
Quantization formats
GGUF
Recommended runtime
Ollama
License
Apache 2.0

Links

More from IFM

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.