DeepSeek-V4-Pro
DeepSeek·
LLMsopen-weightFrontier
Overview
Preview release. 1.6T total / 49B active MoE with hybrid attention and manifold-constrained hyper-connections. Reports 27% of single-token inference FLOPs and 10% of KV cache vs DeepSeek-V3.2 in a 1M-token setting.
Capabilities and innovations
1.6T total / 49B active (MoE)1M context windowThree reasoning effort modes (incl. Think Max)MIT licenseHybrid attention: CSA + HCAManifold-Constrained Hyper-Connections27% inference FLOPs and 10% KV cache vs V3.2 at 1M tokens
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 90.1% | vendor Official model card, instruct table, Think Max column: GPQA Diamond 90.1 (High mode 89.1, Non-Think 72.9). |
| Humanity’s Last Exam (no tools) | 37.7% | vendor Official model card, Think Max column: HLE 37.7. The card does not state whether tools were enabled. |
| SWE-bench Verified | 80.6% | vendor Official model card, Think Max column: SWE-bench Verified 80.6. |
| SWE-bench Pro | 55.4% | vendor Official model card, Think Max column: SWE-bench Pro 55.4. |
| MMLU-Pro | 87.5% | vendor Official model card, Think Max column: MMLU-Pro 87.5. |
| Terminal-Bench 2.x | 67.9% | vendor Terminal-Bench 2.0 67.9 (official model card, Think Max). This is version 2.0, not the 2.1 scores recorded for most 2026 models. |
- MMMU-Pro: not reported by the vendor
Architecture and hardware
- Parameters
- 1600B total, 49B active per token (MoE)
- Estimated VRAM at Q4
- ~920 GB, Frontier class
- Quantization formats
- FP8, GGUF
- Recommended runtime
- vLLM
- License
- MIT
Links
More from DeepSeek
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.