AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

DeepSeek-V4-Pro

DeepSeek·

LLMsopen-weightcloud + localFrontier

Overview

Preview release, available via API and open weights. 1.6T total / 49B active MoE with hybrid attention and manifold-constrained hyper-connections. Reports 27% of single-token inference FLOPs and 10% of KV cache vs DeepSeek-V3.2 in a 1M-token setting.

Capabilities and innovations

1.6T total / 49B active MoE1M token context windowThree reasoning effort modes (incl. Think Max)MIT licenseHybrid attention: CSA + HCAManifold-Constrained Hyper-Connections27% inference FLOPs and 10% KV cache vs V3.2 at 1M tokens

Benchmarks

BenchmarkScoreSource
GPQA Diamond90.1%vendor
Official model card, instruct table, Think Max column: GPQA Diamond 90.1 (High mode 89.1, Non-Think 72.9).
Humanity’s Last Exam (no tools)37.7%vendor
Official model card, Think Max column: HLE 37.7. The card does not state whether tools were enabled.
SWE-bench Verified80.6%vendor
Official model card, Think Max column: SWE-bench Verified 80.6.
SWE-bench Pro55.4%vendor
Official model card, Think Max column: SWE-bench Pro 55.4.
MMLU-Pro87.5%vendor
Official model card, Think Max column: MMLU-Pro 87.5.
Terminal-Bench 2.x67.9%vendor
Terminal-Bench 2.0 67.9 (official model card, Think Max). This is version 2.0, not the 2.1 scores recorded for most 2026 models.
  • MMMU-Pro: not reported by the vendor

API pricing

$0.44 per million input tokens, $0.87 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-07-31)

Architecture and hardware

Parameters
1600B total, 49B active per token (MoE)
Estimated VRAM at Q4
~920 GB, Frontier class

Links

More from DeepSeek

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.