AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

DeepSeek-V4-Flash-Vision-Exp

DeepSeek·

LLMsmultimodalcloud

Overview

Experimental vision variant of DeepSeek-V4-Flash-0731, live on the DeepSeek API platform as model 'deepseek-v4-flash-vision-exp'. Same 284B total / 13B active MoE and 1M-token context as the text build, extended to image input (text output only); images are billed as up to 384 tokens each and can be passed as base64 through Chat Completions, Messages and the Responses API. DeepSeek describes text capability — agents, reasoning, world knowledge — as unchanged versus V4-Flash and reports the gains on agent benchmarks that require visual understanding: Terminal-Bench 2.1 83.9 and DeepSWE 59.3. Unlike the rest of the V4 line this is a closed API release — no open weights, no Hugging Face repository at launch, and the 'Exp' tag marks it as experimental.

Capabilities and innovations

284B total / 13B active MoEImage input, text output1M token context windowMultimodal agent tasksChat Completions, Messages and Responses APIImages billed at up to 384 tokens eachFirst vision-enabled model of the DeepSeek V4 lineVision added on top of the V4-Flash-0731 checkpoint at unchanged parameter countAPI-only release — first V4 build without open weights

Benchmarks

BenchmarkScoreSource
Terminal-Bench 2.x83.9%vendor
Terminal-Bench 2.1 83.9 from the launch announcement (DeepSWE 59.3 in the same table, no schema field). Version 2.1, comparable to the 79.0 recorded for V4-Flash-0731 but not to the 2.0 figures of the April preview. Announcement not fetchable from this environment; value taken from consistent secondary transcriptions and not independently reproduced.
  • GPQA Diamond: not reported by the vendor — DeepSeek states text capability is on par with V4-Flash-0731 but publishes no separate run for this build. The 0731 values are deliberately not carried over — a claim of parity is not a measurement.
  • Humanity’s Last Exam (no tools): not reported by the vendor — Same parity claim as GPQA, no separate run published.
  • SWE-bench Pro: not reported by the vendor — The launch note reports agentic suites (Terminal-Bench 2.1, DeepSWE) only.
  • MMLU-Pro: not reported by the vendor — Same agent-focused launch note.
  • MMMU-Pro: independent evaluation pending — This is the first DeepSeek model where MMMU-Pro would apply, but no multimodal reasoning benchmark was published at launch and no independent run exists yet.

API pricing

$0.22 per million input tokens, $0.66 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-09-01)

Architecture and hardware

Parameters
284B total, 13B active per token (MoE)

Links

More from DeepSeek

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.