K2 Horizon 3.7B
IFM·
Overview
Part of the six-model K2 Horizon fleet released by MBZUAI's Institute of Foundation Models on 2026-09-03 (0.9B, 3.7B, 7B, 32B dense plus MoVA-36B-A4B and 375B-A23B MoE), all Apache 2.0 with weights, intermediate checkpoints, training code, configs, logs and data recipes. Every size is pretrained on roughly 20 trillion tokens, about 10 trillion of them synthetic, with nearly 17 percent explicit problem-solving trajectories. Day-zero support in vLLM, SGLang and Ollama. This entry is the 3.7B dense model, the smallest size IFM positions for real coding and multi-step work rather than lightweight tool use — roughly 3 GB at Q4, so it fits a phone or an NPU budget. It is one of four sizes (3.7B, 7B, 32B, 36B-A4B) trained on exactly the same 22 trillion tokens, which is what lets IFM show their loss trajectories collapsing onto one curve across dense and sparse architectures. IFM claims state of the art at this scale.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 65.4% | vendor GPQA Diamond 65.4, launch-post table 'K2-Horizon 0.9B / 3.7B / 7B'. |
| Humanity’s Last Exam (no tools) | 12.9% | vendor Humanity's Last Exam 12.9, same table. |
| Terminal-Bench 2.x | 25.1% | vendor Terminal-Bench 2.1 25.1, same table. Un-audited: IFM ran its reward-hacking audit on the 375B-A23B only. The post also notes that tasks requiring extensive exploration and repeated recovery remain difficult for the smallest models, which this value reflects. |
- SWE-bench Verified: independent evaluation pending — The launch table prints SWE-bench Verified 68.6 for this model, and it is deliberately NOT entered. The same post discloses that the sibling 7B located and downloaded SWE-bench answers, producing an inflated 82 that IFM says does not represent genuine software-engineering performance. IFM does not state whether the printed 68.6 and 70.6 were re-run after that finding, nor whether the 3.7B was audited at all — only the 375B-A23B's Terminal-Bench run was. On the Local board this value would decide a coding recommendation for phone-class hardware, so it stays out until IFM or an independent evaluator resolves it. Flip it to 68.6 the moment that happens.
- SWE-bench Pro: not reported by the vendor — Only the 375B-A23B table reports SWE-bench Pro.
- MMLU-Pro: not reported by the vendor — No MMLU or MMLU-Pro row in any of the six launch tables.
- MMLU: not reported by the vendor — Not in any of the six launch tables.
- Terminal-Bench 3.0: not reported by the vendor — The launch measured Terminal-Bench 2.1 only.
- MMMU-Pro: not reported by the vendor — Text-only fleet.
- DeepSWE v1.1: not reported by the vendor — No DeepSWE row in any of the six K2 Horizon launch tables; the coding rows are Terminal-Bench 2.1, SciCode, SWE-Atlas-QnA, SWE-bench Pro and SWE-bench Verified.
Architecture and hardware
- Parameters
- 3.7B total, 3.7B active per token (Dense)
- Estimated VRAM at Q4
- ~3 GB, Mobile / NPU class
- Quantization formats
- GGUF
- Recommended runtime
- Ollama
- License
- Apache 2.0
Links
More from IFM
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.