AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

DeepSeek-V4.1-Flash

DeepSeek·

LLMsopen-weightcloud + localFrontier

Overview

Natively multimodal successor to DeepSeek-V4-Flash-0731 and the first model of DeepSeek's Causal Encoder-Decoder family: 40 layers split into a 20-layer causal encoder and a 20-layer decoder, a 552B backbone plus 196B of sparsely-accessed Engram memory, activating 8B parameters per token during prefill and 16B during decode. 1M-token context, MIT license, open weights. FP4 main KV caching at 890 bytes/token and DSpark speculative decoding drive a reported 409.5 tokens/s end-to-end, and DeepSeek prices it at roughly a quarter of V4-Pro.

Capabilities and innovations

552B backbone MoE, 8B active (prefill) / 16B active (decode)Native multimodal image and text input1M token context window384 routed experts (1 shared, 6 routed per token)MIT license, open weightsCausal Encoder-Decoder (CED) architectureCompressed Sparse Attention 2 with hierarchical sparse indexer196B sparsely-accessed Engram memoryFP4 main KV cache at 890 bytes/tokenDSpark speculative decodingSingle-Pass mHC

Benchmarks

BenchmarkScoreSource
GPQA Diamond90.9%vendor
GPQA Diamond 90.9, DeepSeek launch table. No independent run yet: Artificial Analysis, which supplied the 0731 entry's GPQA, has not published a V4.1 evaluation.
Humanity’s Last Exam (no tools)36.8%vendor
HLE 36.8, DeepSeek launch table. DeepSeek also prints 39.1 on the pure-text subset; the full-set figure is used, since that is what every other HLE value in the catalog measures.
Terminal-Bench 2.x90.6%vendor
Terminal-Bench 2.1 90.6, DeepSeek self-report, run in the Minimal mode of DeepSeek's own harness at 1M context (stated on the model card) rather than in the harness the earlier 2.x values were measured under. Two reasons to hold it loosely until an independent run lands: it is the top 2.x value in the catalog, and on the predecessor DeepSeek self-reported 82.7 where Artificial Analysis measured 79.0, so this vendor line ran 3.7 points hot the last time both existed. Version 2.1, not comparable with the 2.0 scores on earlier entries.
Terminal-Bench 3.030%vendor
Terminal-Bench 3.0 30.0, DeepSeek self-report under its own harness. Its own generation on its own scale, not comparable with the 2.x or 4.0 fields.
Terminal-Bench 4.031.2%vendor
Terminal-Bench 4.0 31.2, DeepSeek self-report. First 4.0 value in the catalog from a Chinese lab and from an open-weight model, which is the coverage condition the field's FRONTIER_BENCHMARKS entry was held open on. It is not flipped in the same change: unlike the seven existing 4.0 values, which cross-check against each other through the Harbor harness, this one is explicitly run in DeepSeek's own harness, so the bloc condition is met while comparability is newly in question.
DeepSWE v1.174.2%vendor
DeepSWE v1.1 74.2, DeepSeek self-report, explicitly labelled v1.1 (DeepSeek's earlier unversioned 59.3 stays out of the field). Measured in DeepSeek's own harness, where DeepSWE fixes runs to mini-swe-agent, so this carries the same harness caveat the Muse Spark entries carry. The 74.0 and 73.0 DeepSeek prints for Claude Opus 5 and GPT-5.6 Sol are DeepSeek measuring other labs' models and are not entered on those entries.
  • SWE-bench Verified: not reported by the vendor
  • SWE-bench Pro: not reported by the vendor — DeepSeek reports DeepSWE v1.1 in its place, the same substitution nine catalog entries since Jun 2026 record.
  • MMLU-Pro: independent evaluation pending — DeepSeek publishes no MMLU-Pro value for this build. The 0731 entry's MMLU-Pro came from Vals AI, which has not yet run V4.1. Without it this entry carries only 2 of the 4 era values (GPQA and Terminal-Bench 2.x) and stays out of the 2026+ LLM leadership race; it does enter the coding race, which needs only one coding signal.
  • MMMU-Pro: not reported by the vendor — No multimodal reasoning benchmark published despite native image input.

API pricing

$0.15 per million input tokens, $0.6 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-09-10)

Architecture and hardware

Parameters
552B total, 16B active per token (MoE)
Estimated VRAM at Q4
~318 GB, Frontier class

Links

More from DeepSeek

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.