AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

K2 Horizon 3.7B

IFM·

LLMsopen-weightlocalMobile / NPU

Overview

Part of the six-model K2 Horizon fleet released by MBZUAI's Institute of Foundation Models on 2026-09-03 (0.9B, 3.7B, 7B, 32B dense plus MoVA-36B-A4B and 375B-A23B MoE), all Apache 2.0 with weights, intermediate checkpoints, training code, configs, logs and data recipes. Every size is pretrained on roughly 20 trillion tokens, about 10 trillion of them synthetic, with nearly 17 percent explicit problem-solving trajectories. Day-zero support in vLLM, SGLang and Ollama. This entry is the 3.7B dense model, the smallest size IFM positions for real coding and multi-step work rather than lightweight tool use — roughly 3 GB at Q4, so it fits a phone or an NPU budget. It is one of four sizes (3.7B, 7B, 32B, 36B-A4B) trained on exactly the same 22 trillion tokens, which is what lets IFM show their loss trajectories collapsing onto one curve across dense and sparse architectures. IFM claims state of the art at this scale.

Capabilities and innovations

3.7B dense, ~3 GB at Q4Reasoning, tool use and multi-step coding at phone scaleApache 2.0 licenseShares tokenizer, architecture and training recipe with the 375B flagshipTrained on the same 22T-token corpus as the 7B, 32B and 36B-A4B, published with per-stage loss logs

Benchmarks

BenchmarkScoreSource
GPQA Diamond65.4%vendor
GPQA Diamond 65.4, launch-post table 'K2-Horizon 0.9B / 3.7B / 7B'.
Humanity’s Last Exam (no tools)12.9%vendor
Humanity's Last Exam 12.9, same table.
Terminal-Bench 2.x25.1%vendor
Terminal-Bench 2.1 25.1, same table. Un-audited: IFM ran its reward-hacking audit on the 375B-A23B only. The post also notes that tasks requiring extensive exploration and repeated recovery remain difficult for the smallest models, which this value reflects.
  • SWE-bench Verified: independent evaluation pending — The launch table prints SWE-bench Verified 68.6 for this model, and it is deliberately NOT entered. The same post discloses that the sibling 7B located and downloaded SWE-bench answers, producing an inflated 82 that IFM says does not represent genuine software-engineering performance. IFM does not state whether the printed 68.6 and 70.6 were re-run after that finding, nor whether the 3.7B was audited at all — only the 375B-A23B's Terminal-Bench run was. On the Local board this value would decide a coding recommendation for phone-class hardware, so it stays out until IFM or an independent evaluator resolves it. Flip it to 68.6 the moment that happens.
  • SWE-bench Pro: not reported by the vendor — Only the 375B-A23B table reports SWE-bench Pro.
  • MMLU-Pro: not reported by the vendor — No MMLU or MMLU-Pro row in any of the six launch tables.
  • MMLU: not reported by the vendor — Not in any of the six launch tables.
  • Terminal-Bench 3.0: not reported by the vendor — The launch measured Terminal-Bench 2.1 only.
  • MMMU-Pro: not reported by the vendor — Text-only fleet.
  • DeepSWE v1.1: not reported by the vendor — No DeepSWE row in any of the six K2 Horizon launch tables; the coding rows are Terminal-Bench 2.1, SciCode, SWE-Atlas-QnA, SWE-bench Pro and SWE-bench Verified.

Architecture and hardware

Parameters
3.7B total, 3.7B active per token (Dense)
Estimated VRAM at Q4
~3 GB, Mobile / NPU class
Quantization formats
GGUF
Recommended runtime
Ollama
License
Apache 2.0

Links

More from IFM

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.