DeepSeek-V4-Flash
DeepSeek·
LLMsopen-weightMulti-GPU Self-Host
Overview
Preview release, superseded by DeepSeek-V4-Flash-0731 on 2026-07-31. Smaller 284B total / 13B active MoE sibling of V4-Pro, sharing hybrid attention and hyper-connection architecture.
Capabilities and innovations
284B total / 13B active (MoE)1M context windowThree reasoning effort modes (incl. Think Max)MIT licenseHybrid attention: CSA + HCAManifold-Constrained Hyper-ConnectionsShared architecture with V4-Pro at smaller scale
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 88.1% | vendor Official model card, instruct table, Think Max column: GPQA Diamond 88.1. |
| Humanity’s Last Exam (no tools) | 34.8% | vendor Official model card, Think Max column: HLE 34.8. The card does not state whether tools were enabled. |
| SWE-bench Verified | 79% | vendor Official model card, Think Max column: SWE-bench Verified 79.0. |
| SWE-bench Pro | 52.6% | vendor Official model card, Think Max column: SWE-bench Pro 52.6. |
| MMLU-Pro | 86.2% | vendor Official model card, Think Max column: MMLU-Pro 86.2. |
| Terminal-Bench 2.x | 56.9% | vendor Terminal-Bench 2.0 56.9 (official model card). This is version 2.0, not the 2.1 scores recorded for most 2026 models. |
- MMMU-Pro: not reported by the vendor
Architecture and hardware
- Parameters
- 284B total, 13B active per token (MoE)
- Estimated VRAM at Q4
- ~164 GB, Multi-GPU Self-Host class
- Quantization formats
- FP8, GGUF
- Recommended runtime
- vLLM
- License
- MIT
Links
More from DeepSeek
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.