AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

DeepSeek-V4-Pro

DeepSeek·

LLMsopen-weightFrontier

Overview

Preview release. 1.6T total / 49B active MoE with hybrid attention and manifold-constrained hyper-connections. Reports 27% of single-token inference FLOPs and 10% of KV cache vs DeepSeek-V3.2 in a 1M-token setting.

Capabilities and innovations

1.6T total / 49B active (MoE)1M context windowThree reasoning effort modes (incl. Think Max)MIT licenseHybrid attention: CSA + HCAManifold-Constrained Hyper-Connections27% inference FLOPs and 10% KV cache vs V3.2 at 1M tokens

Benchmarks

BenchmarkScoreSource
GPQA Diamond90.1%vendor
Official model card, instruct table, Think Max column: GPQA Diamond 90.1 (High mode 89.1, Non-Think 72.9).
Humanity’s Last Exam (no tools)37.7%vendor
Official model card, Think Max column: HLE 37.7. The card does not state whether tools were enabled.
SWE-bench Verified80.6%vendor
Official model card, Think Max column: SWE-bench Verified 80.6.
SWE-bench Pro55.4%vendor
Official model card, Think Max column: SWE-bench Pro 55.4.
MMLU-Pro87.5%vendor
Official model card, Think Max column: MMLU-Pro 87.5.
Terminal-Bench 2.x67.9%vendor
Terminal-Bench 2.0 67.9 (official model card, Think Max). This is version 2.0, not the 2.1 scores recorded for most 2026 models.
  • MMMU-Pro: not reported by the vendor

Architecture and hardware

Parameters
1600B total, 49B active per token (MoE)
Estimated VRAM at Q4
~920 GB, Frontier class
Quantization formats
FP8, GGUF
Recommended runtime
vLLM
License
MIT

Links

More from DeepSeek

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.