AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

DeepSeek-V4-Flash-0731

DeepSeek·

LLMsopen-weightMulti-GPU Self-Host

Overview

General-availability build of DeepSeek-V4-Flash. Identical 284B total / 13B active MoE architecture and size as the April preview — re-post-trained, with the gains concentrated in agentic tool use. 1M-token context, MIT license. On 2026-08-21 DeepSeek added a vision variant on top of this checkpoint, DeepSeek-V4-Flash-Vision-Exp — API-only, without open weights.

Capabilities and innovations

284B total / 13B active MoE1M token context windowThree reasoning effort modes (incl. Think Max)Agentic tool use (Terminal-Bench, NL2Repo, Cybergym, Toolathlon)MIT licenseHybrid attention: CSA + HCAManifold-Constrained Hyper-ConnectionsAgent-focused re-post-training on the preview checkpoint

Benchmarks

BenchmarkScoreSource
GPQA Diamond91%independent
GPQA Diamond 91 — Artificial Analysis, measured on the 0731 build (+1 point over the April V4 Flash).
Humanity’s Last Exam (no tools)37%independent
HLE 37 no-tools — Artificial Analysis Intelligence Index component (+5 points over the April V4 Flash).
Terminal-Bench 2.x79%independent
Terminal-Bench 2.1 79.0 — Artificial Analysis (+17 points over the April V4 Flash). DeepSeek self-reports 82.7 on the same benchmark version; the independent value is used as primary. Version 2.1, not comparable to the 2.0 scores recorded for earlier models.
  • SWE-bench Verified: independent evaluation pending — The 0731 changelog reports only agentic suites. The preview checkpoint's 79.0 SWE-bench Verified is not carried over.
  • SWE-bench Pro: independent evaluation pending — Preview checkpoint scored 52.6; no re-run published for 0731.
  • MMLU-Pro: independent evaluation pending — Preview checkpoint scored 86.2; no re-run published for 0731.
  • MMMU-Pro: not reported by the vendor

Architecture and hardware

Parameters
284B total, 13B active per token (MoE)
Estimated VRAM at Q4
~164 GB, Multi-GPU Self-Host class
Quantization formats
FP8, GGUF
Recommended runtime
vLLM
License
MIT

Links

More from DeepSeek

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.