AI Model Timeline

Enterprise and RAG suitability, measured independently: grounded-summarization hallucination rate, multi-turn agentic tool use and knowledge-grounded agentic workflows.

Live Status: Data through 2026-09-10
Reliability

Enterprise & RAG Leaderboard

The axis academic benchmarks miss: how dependable a model is in retrieval-augmented and agentic deployments. Three axes, all oriented higher = better — hallucination is shown inverted as Grounding so the table never reads backwards. They come from different evaluations and are not a combined score: read each column on its own. Only curated flagships are scored; marks a genuine gap, not a zero.

Newest scored model: Sep 2026·Grounding: no new snapshot (May 2026)Tool-Use: superseded upstream (Jun 2026)τ³-Banking: publishing (Sep 2026)
#Model
Grounding 100 − hallucination
Tool-Use τ-bench
τ³-Banking policy KB + tool calls
1
Qwen3.8-MaxCN
51.3
2
Grok 4.6US
50.7
3
GLM-5.3CN
50.3
4
Qwen3.8-2.4T-A95BCN
49.1
5
Claude Fable 5.1US
47.2
6
Muse Spark 1.3US
47.0
7
Kimi K3CN
46.0
8
GPT-5.6 SolUS
44.3
9
Claude Opus 5US
42.1
10
DeepSeek-V4-Pro-0813CN
39.6
11
Claude Fable 5US
89.2
38.1
12
ChatGPT 5.1US
87.9
74.2
13
GPT-5.3 CodexUS
77.8
14
Gemini 3US
86.4
75.3
15
Gemini 3.1 ProUS
89.6
76.5
16
Claude Opus 4.5US
89.1
17
Claude Opus 4.6US
87.8
84.8
18
Mistral 3 FamilyEU
85.5
70.2
19
Grok 4.1US
80.8
74.6
20
DeepSeek-V3.2CN
93.7
71.3
21
DeepSeek-V4-ProCN
91.4
22
Qwen3.5CN
89.3
77.5
23
GPT-5.4 ThinkingUS
93.0
24
GPT-5.4 ProUS
91.7
80.1
25
Kimi K2.5CN
85.8
74.2
26
GLM-5CN
89.9
82.1
27
Claude Opus 4.7US
88.0
28
GPT-5.5US
90.7
29
Qwen3.6 PlusCN
76.8
30
Command A+CA
86.0*

30 flagships scored · coverage: Grounding 17/30 · Tool-Use 14/30 · τ³-Banking 11/30. Click a column to sort.

Grounding ↑

Inverted Vectara HHEM hallucination rate (100 − rate). HHEM scores faithfulness when summarizing a given document — a proxy for RAG grounding, not full open-domain factuality.

Tool-Use ↑

τ-bench multi-turn task completion with policy adherence (retail/airline). Frozen: the authority stopped updating the task set. Kept because it is the only tool-use figure the 2025 and early-2026 models carry.

τ³-Banking ↑

τ-bench’s successor family, banking domain, run by Artificial Analysis. Find the right policy in ~700 linked documents, then execute the tool calls it prescribes. The only axis here that scores 2026-H2 models — and far harder than the τ-bench column, so never read the two side by side.

Why recent models are missing
  • Grounding · Vectara HHEM (grounded summarization) — newest published snapshot is 2026-05-11. Models released after it have no figure on this axis yet; the gap is upstream, not a missing entry.
  • Function-Calling · Berkeley BFCL v4 (overall accuracy) — column not shown. Column is not rendered: ACTIVE_COLUMNS drops any metric with this status. BFCL has scored almost none of the frontier since 2025, and standalone call accuracy overlaps with the Tool-Use axis. Kept collectable so a revival needs no backfill — but it is not advertised anywhere the reader can see, and the page/SEO copy must not promise it.
  • Tool-Use · τ-bench (tau1 retail+airline aggregate) — the authority has moved on to τ³-Banking, carried by the `tau3_banking` column. Versions score different task sets, so their numbers are not interchangeable: adopting the successor means a new column here, never an overwrite of these values.

Enterprise reliability is measured by a handful of small teams, and their leaderboards stop and restart. A blank recent row here means nobody independent has scored the model — which is itself worth knowing before shipping it into a RAG pipeline.

The two axes come from different evaluations with different methodologies and snapshot dates — read them as directional signals per axis, not a single combined ranking. Hover any score for its source.

* = value from an alternative benchmark than the column's canonical source (hover for which) — e.g. Command A+ has no Vectara HHEM entry, so its Grounding uses AA-Omniscience Non-Hallucination instead.

Origin badges (US · CN · EU · CA) make the geopolitics explicit: scored here are 17 US, 11 China, 1 EU, 1 Canada. The frontier is a US–China race at the top, while the enterprise/RAG axis is exactly where the EU/Canada players (Mistral, Cohere) compete — when the canonical boards list them.