The benchmark frontier over time: which model held the best reported score on each benchmark, and when it was overtaken.
Cloud Models Local Models The Model Race Reliability Live Status: Data through 2026-09-01
The Model Race
Benchmark Frontier Closed · Cloud vs Open · Local US / West vs China
GPQA Diamond SWE-bench Verified MMLU Pro Humanity's Last Exam SWE-bench Pro Terminal-Bench 2.1 Terminal-Bench 3.0 ARC-AGI-1 ARC-AGI-2
Step Smooth ⊟ Models shown
Closed-weight(Cloud Models) 95.2% Open-weight(Local Models) 93.54% 20 40 60 80 100 2023 2024 2025 2026 GPT-5.6 Sol GPT-5.4 Pro Gemini 3.1 Pro Qwen3.8-2.4T-A95B Kimi K3 Claude Opus 4.5 Gemini 3 Hy3 DeepSeek-V4-Pro Grok 4.1 Claude Opus 4 Qwen3-Max-Thinking Claude Sonnet 4 Claude 3.7 Sonnet DeepSeek-Terminus DeepSeek-V3.1 Grok-3 o3-mini OpenAI o1 GPT-OSS DeepSeek-R1 Claude 3.5 Sonnet DeepSeek-V3 Gemini 1.5 Pro Mistral Large 2 Llama 3.1 405B DeepSeek-V2 Mistral Large GPT-4 Turbo GPT-4 ChatGPT Open −1.7 pt Closed = Cloud Models · Open = Local Models. Each line is the cumulative best license-class score on GPQA Diamond — it rises only when a model sets a new record for its open vs closed bloc. The shrinking vertical gap is the catch-up story. Benchmarks rotate as they saturate, so each shows only the span it covers.
© 2026 AI Model Timeline. Tracking 7 years of progress.