AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

DeepSeek-V4-Pro-0813

DeepSeek·

LLMsopen-weightcloud + localFrontier

Overview

General-availability build of DeepSeek-V4-Pro, replacing the April preview. Same 1.6T total / 49B active MoE architecture as the preview — re-post-trained, with the gains concentrated in agentic and software-engineering tasks. The deepseek-v4-pro API endpoint now points to this build. 1M-token context, MIT license.

Capabilities and innovations

1.6T total / 49B active MoE1M token context windowThree reasoning effort modes (incl. Think Max)MIT licenseHybrid attention: CSA + HCAManifold-Constrained Hyper-ConnectionsRe-post-training of the April preview checkpoint

Benchmarks

BenchmarkScoreSource
GPQA Diamond92.83%third party
GPQA Diamond 92.83%, Artificial Analysis, reasoning max effort, measured on the 0813 build (April preview: 89.4 on vals.ai, 90.1 vendor). Same source lineage as the V4-Flash-0731 entry.
Humanity’s Last Exam (no tools)39.34%third party
HLE 39.34% no-tools, Artificial Analysis Intelligence Index component, max effort (April preview: 37.7 vendor).
SWE-bench Verified96.4%independent
SWE-bench Verified 96.40 (± 0.83), vals.ai, max effort (April preview: 77.4 on the same harness). Era-1 benchmark, near saturation — recorded for continuity, not ranked.
MMLU-Pro86.97%independent
MMLU Pro 86.97%, measured by Vals AI on the 0813 build (accessed 2026-08-17); the April preview scored 87.25 on the same harness. Vals AI measures GPQA Diamond at 92.42 on this run, 0.4 points below the Artificial Analysis value kept as primary.
Terminal-Bench 2.x78.65%third party
Terminal-Bench 2.1 78.65%, Artificial Analysis, max effort. Measurements diverge widely on this build: vals.ai reports 54.68 (± 1.50) and DeepSeek self-reports 87.9 on the same benchmark version. Artificial Analysis is kept as primary for continuity with the V4-Flash-0731 entry. Version 2.1, not comparable to the 2.0 score recorded for the April preview.
  • SWE-bench Pro: independent evaluation pending — The April preview checkpoint scored 55.4; no re-run published for the 0813 build.
  • MMMU-Pro: not reported by the vendor

API pricing

$0.435 per million input tokens, $0.87 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-08-13)

Architecture and hardware

Parameters
1600B total, 49B active per token (MoE)
Estimated VRAM at Q4
~920 GB, Frontier class

Links

More from DeepSeek

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.