AI Model Timeline

The benchmark frontier over time: which model held the best reported score on each benchmark, and when it was overtaken.

Live Status: Data through 2026-09-01
The Model Race

Benchmark Frontier

Closed-weight(Cloud Models)95.2%Open-weight(Local Models)93.54%
204060801002023202420252026GPT-5.6 SolGPT-5.4 ProGemini 3.1 ProQwen3.8-2.4T-A95BKimi K3Claude Opus 4.5Gemini 3Hy3DeepSeek-V4-ProGrok 4.1Claude Opus 4Qwen3-Max-ThinkingClaude Sonnet 4Claude 3.7 SonnetDeepSeek-TerminusDeepSeek-V3.1Grok-3o3-miniOpenAI o1GPT-OSSDeepSeek-R1Claude 3.5 SonnetDeepSeek-V3Gemini 1.5 ProMistral Large 2Llama 3.1 405BDeepSeek-V2Mistral LargeGPT-4 TurboGPT-4ChatGPTOpen −1.7 pt
Closed = Cloud Models · Open = Local Models. Each line is the cumulative best license-class score on GPQA Diamond — it rises only when a model sets a new record for its open vs closed bloc. The shrinking vertical gap is the catch-up story. Benchmarks rotate as they saturate, so each shows only the span it covers.