AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

DeepSeek-V3

DeepSeek·

LLMsopen-weightFrontier

Overview

Cost-efficient SOTA model using FP8 mixed-precision training. Achieved GPT-4o-level performance at a fraction of the training cost ($5.5M).

Capabilities and innovations

671B total / 37B active parameters (MoE)128K context windowMulti-Token PredictionStrong math and codingFP8 mixed-precision trainingMulti-Token Prediction (MTP)Auxiliary-loss-free load balancingDualPipe pipeline parallelism

Benchmarks

BenchmarkScoreSource
GPQA Diamond59.1%unsourced
SWE-bench Verified42%unsourced
MMLU87.1%unsourced

Architecture and hardware

Parameters
671B total, 37B active per token (MoE)
Estimated VRAM at Q4
~386 GB, Frontier class
Quantization formats
GGUF, GPTQ, FP8
Recommended runtime
vLLM
License
MIT

Links

More from DeepSeek

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.