AI Model Timeline

Enterprise and RAG suitability, measured independently: grounded-summarization hallucination rate, function-calling accuracy and multi-turn tool use.

Live Status: Data through 2026-09-01
Reliability

Enterprise & RAG Leaderboard

The axis academic benchmarks miss: how dependable a model is in retrieval-augmented and agentic deployments. Two current-state axes, both oriented higher = better — hallucination is shown inverted as Grounding so the table never reads backwards. Only curated current flagships are scored; marks a genuine gap, not a zero.

#Model
Grounding 100 − hallucination
Tool-Use τ-bench
1
DeepSeek-V3.2CN
93.7
71.3
2
GPT-5.4 ProUS
91.7
80.1
3
GPT-5.5US
90.7
4
GLM-5CN
89.9
82.1
5
Gemini 3.1 ProUS
89.6
76.5
6
Qwen3.5CN
89.3
77.5
7
Claude Opus 4.5US
89.1
8
Claude Opus 4.7US
88.0
9
ChatGPT 5.1US
87.9
74.2
10
Claude Opus 4.6US
87.8
84.8
11
Gemini 3US
86.4
75.3
12
Command A+CA
86.0*
13
Kimi K2.5CN
85.8
74.2
14
Mistral 3 FamilyEU
85.5
70.2
15
Grok 4.1US
80.8
74.6
16
GPT-5.3 CodexUS
77.8
17
Qwen3.6 PlusCN
76.8
18
Claude Fable 5US
89.2

18 flagships scored · coverage: Grounding 15/18 · Tool-Use 14/18. Click a column to sort. Function-Calling (BFCL v4) is collected (2 so far) but withheld until the benchmark covers more of the frontier.

Grounding ↑

Inverted Vectara HHEM hallucination rate (100 − rate). HHEM scores faithfulness when summarizing a given document — a proxy for RAG grounding, not full open-domain factuality.

Tool-Use ↑

τ-bench multi-turn task completion with policy adherence (retail/airline). The hardest, sparsest metric — many models have no published score yet.

The two axes come from different evaluations with different methodologies and snapshot dates — read them as directional signals per axis, not a single combined ranking. Hover any score for its source.

* = value from an alternative benchmark than the column's canonical source (hover for which) — e.g. Command A+ has no Vectara HHEM entry, so its Grounding uses AA-Omniscience Non-Hallucination instead.

Origin badges (US · CN · EU · CA) make the geopolitics explicit: scored here are 11 US, 5 China, 1 EU, 1 Canada. The frontier is a US–China race at the top, while the enterprise/RAG axis is exactly where the EU/Canada players (Mistral, Cohere) compete — when the canonical boards list them.