AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

DeepSeek-V4-Flash

DeepSeek·

LLMsopen-weightcloud + localMulti-GPU Self-Host

Overview

Preview release, available via API and open weights; superseded by DeepSeek-V4-Flash-0731 on 2026-07-31. Smaller 284B total / 13B active MoE sibling of V4-Pro, sharing hybrid attention and hyper-connection architecture.

Capabilities and innovations

284B total / 13B active MoE1M token context windowThree reasoning effort modes (incl. Think Max)MIT licenseHybrid attention: CSA + HCAManifold-Constrained Hyper-ConnectionsShared architecture with V4-Pro at smaller scale

Benchmarks

BenchmarkScoreSource
GPQA Diamond88.1%vendor
Official model card, instruct table, Think Max column: GPQA Diamond 88.1.
Humanity’s Last Exam (no tools)34.8%vendor
Official model card, Think Max column: HLE 34.8. The card does not state whether tools were enabled.
SWE-bench Verified79%vendor
Official model card, Think Max column: SWE-bench Verified 79.0.
SWE-bench Pro52.6%vendor
Official model card, Think Max column: SWE-bench Pro 52.6.
MMLU-Pro86.2%vendor
Official model card, Think Max column: MMLU-Pro 86.2.
Terminal-Bench 2.x56.9%vendor
Terminal-Bench 2.0 56.9 (official model card, Think Max). This is version 2.0, not the 2.1 scores recorded for most 2026 models.
  • MMMU-Pro: not reported by the vendor

API pricing

$0.08092 per million input tokens, $0.16184 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-09-01)

  • 2026-07-31: $0.14 / $0.28 per MTok
  • 2026-07-31: $0.14 / $0.28 per MTok

Architecture and hardware

Parameters
284B total, 13B active per token (MoE)
Estimated VRAM at Q4
~164 GB, Multi-GPU Self-Host class

Links

More from DeepSeek

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.