AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

DeepSeek-V4-Flash

DeepSeek·

LLMsopen-weightMulti-GPU Self-Host

Overview

Preview release, superseded by DeepSeek-V4-Flash-0731 on 2026-07-31. Smaller 284B total / 13B active MoE sibling of V4-Pro, sharing hybrid attention and hyper-connection architecture.

Capabilities and innovations

284B total / 13B active (MoE)1M context windowThree reasoning effort modes (incl. Think Max)MIT licenseHybrid attention: CSA + HCAManifold-Constrained Hyper-ConnectionsShared architecture with V4-Pro at smaller scale

Benchmarks

BenchmarkScoreSource
GPQA Diamond88.1%vendor
Official model card, instruct table, Think Max column: GPQA Diamond 88.1.
Humanity’s Last Exam (no tools)34.8%vendor
Official model card, Think Max column: HLE 34.8. The card does not state whether tools were enabled.
SWE-bench Verified79%vendor
Official model card, Think Max column: SWE-bench Verified 79.0.
SWE-bench Pro52.6%vendor
Official model card, Think Max column: SWE-bench Pro 52.6.
MMLU-Pro86.2%vendor
Official model card, Think Max column: MMLU-Pro 86.2.
Terminal-Bench 2.x56.9%vendor
Terminal-Bench 2.0 56.9 (official model card). This is version 2.0, not the 2.1 scores recorded for most 2026 models.
  • MMMU-Pro: not reported by the vendor

Architecture and hardware

Parameters
284B total, 13B active per token (MoE)
Estimated VRAM at Q4
~164 GB, Multi-GPU Self-Host class
Quantization formats
FP8, GGUF
Recommended runtime
vLLM
License
MIT

Links

More from DeepSeek

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.