Enterprise & RAG Leaderboard
The axis academic benchmarks miss: how dependable a model is in retrieval-augmented and agentic deployments. Three axes, all oriented higher = better — hallucination is shown inverted as Grounding so the table never reads backwards. They come from different evaluations and are not a combined score: read each column on its own. Only curated flagships are scored; — marks a genuine gap, not a zero.
| # | Model | Grounding ↑↕100 − hallucination | Tool-Use ↑↕τ-bench | τ³-Banking ↑▼policy KB + tool calls |
|---|---|---|---|---|
| 1 | Qwen3.8-MaxCN | — | — | 51.3 |
| 2 | Grok 4.6US | — | — | 50.7 |
| 3 | GLM-5.3CN | — | — | 50.3 |
| 4 | Qwen3.8-2.4T-A95BCN | — | — | 49.1 |
| 5 | Claude Fable 5.1US | — | — | 47.2 |
| 6 | Muse Spark 1.3US | — | — | 47.0 |
| 7 | Kimi K3CN | — | — | 46.0 |
| 8 | GPT-5.6 SolUS | — | — | 44.3 |
| 9 | Claude Opus 5US | — | — | 42.1 |
| 10 | DeepSeek-V4-Pro-0813CN | — | — | 39.6 |
| 11 | Claude Fable 5US | — | 89.2 | 38.1 |
| 12 | ChatGPT 5.1US | 87.9 | 74.2 | — |
| 13 | GPT-5.3 CodexUS | — | 77.8 | — |
| 14 | Gemini 3US | 86.4 | 75.3 | — |
| 15 | Gemini 3.1 ProUS | 89.6 | 76.5 | — |
| 16 | Claude Opus 4.5US | 89.1 | — | — |
| 17 | Claude Opus 4.6US | 87.8 | 84.8 | — |
| 18 | Mistral 3 FamilyEU | 85.5 | 70.2 | — |
| 19 | Grok 4.1US | 80.8 | 74.6 | — |
| 20 | DeepSeek-V3.2CN | 93.7 | 71.3 | — |
| 21 | DeepSeek-V4-ProCN | 91.4 | — | — |
| 22 | Qwen3.5CN | 89.3 | 77.5 | — |
| 23 | GPT-5.4 ThinkingUS | 93.0 | — | — |
| 24 | GPT-5.4 ProUS | 91.7 | 80.1 | — |
| 25 | Kimi K2.5CN | 85.8 | 74.2 | — |
| 26 | GLM-5CN | 89.9 | 82.1 | — |
| 27 | Claude Opus 4.7US | 88.0 | — | — |
| 28 | GPT-5.5US | 90.7 | — | — |
| 29 | Qwen3.6 PlusCN | — | 76.8 | — |
| 30 | Command A+CA | 86.0* | — | — |
30 flagships scored · coverage: Grounding 17/30 · Tool-Use 14/30 · τ³-Banking 11/30. Click a column to sort.
Inverted Vectara HHEM hallucination rate (100 − rate). HHEM scores faithfulness when summarizing a given document — a proxy for RAG grounding, not full open-domain factuality.
τ-bench multi-turn task completion with policy adherence (retail/airline). Frozen: the authority stopped updating the task set. Kept because it is the only tool-use figure the 2025 and early-2026 models carry.
τ-bench’s successor family, banking domain, run by Artificial Analysis. Find the right policy in ~700 linked documents, then execute the tool calls it prescribes. The only axis here that scores 2026-H2 models — and far harder than the τ-bench column, so never read the two side by side.
- Grounding · Vectara HHEM (grounded summarization) — newest published snapshot is 2026-05-11. Models released after it have no figure on this axis yet; the gap is upstream, not a missing entry.
- Function-Calling · Berkeley BFCL v4 (overall accuracy) — column not shown. Column is not rendered: ACTIVE_COLUMNS drops any metric with this status. BFCL has scored almost none of the frontier since 2025, and standalone call accuracy overlaps with the Tool-Use axis. Kept collectable so a revival needs no backfill — but it is not advertised anywhere the reader can see, and the page/SEO copy must not promise it.
- Tool-Use · τ-bench (tau1 retail+airline aggregate) — the authority has moved on to τ³-Banking, carried by the `tau3_banking` column. Versions score different task sets, so their numbers are not interchangeable: adopting the successor means a new column here, never an overwrite of these values.
Enterprise reliability is measured by a handful of small teams, and their leaderboards stop and restart. A blank recent row here means nobody independent has scored the model — which is itself worth knowing before shipping it into a RAG pipeline.
The two axes come from different evaluations with different methodologies and snapshot dates — read them as directional signals per axis, not a single combined ranking. Hover any score for its source.
* = value from an alternative benchmark than the column's canonical source (hover for which) — e.g. Command A+ has no Vectara HHEM entry, so its Grounding uses AA-Omniscience Non-Hallucination instead.
Origin badges (US · CN · EU · CA) make the geopolitics explicit: scored here are 17 US, 11 China, 1 EU, 1 Canada. The frontier is a US–China race at the top, while the enterprise/RAG axis is exactly where the EU/Canada players (Mistral, Cohere) compete — when the canonical boards list them.