Enterprise & RAG Leaderboard
The axis academic benchmarks miss: how dependable a model is in retrieval-augmented and agentic deployments. Two current-state axes, both oriented higher = better — hallucination is shown inverted as Grounding so the table never reads backwards. Only curated current flagships are scored; — marks a genuine gap, not a zero.
| # | Model | Grounding ↑▼100 − hallucination | Tool-Use ↑↕τ-bench |
|---|---|---|---|
| 1 | DeepSeek-V3.2CN | 93.7 | 71.3 |
| 2 | GPT-5.4 ProUS | 91.7 | 80.1 |
| 3 | GPT-5.5US | 90.7 | — |
| 4 | GLM-5CN | 89.9 | 82.1 |
| 5 | Gemini 3.1 ProUS | 89.6 | 76.5 |
| 6 | Qwen3.5CN | 89.3 | 77.5 |
| 7 | Claude Opus 4.5US | 89.1 | — |
| 8 | Claude Opus 4.7US | 88.0 | — |
| 9 | ChatGPT 5.1US | 87.9 | 74.2 |
| 10 | Claude Opus 4.6US | 87.8 | 84.8 |
| 11 | Gemini 3US | 86.4 | 75.3 |
| 12 | Command A+CA | 86.0* | — |
| 13 | Kimi K2.5CN | 85.8 | 74.2 |
| 14 | Mistral 3 FamilyEU | 85.5 | 70.2 |
| 15 | Grok 4.1US | 80.8 | 74.6 |
| 16 | GPT-5.3 CodexUS | — | 77.8 |
| 17 | Qwen3.6 PlusCN | — | 76.8 |
| 18 | Claude Fable 5US | — | 89.2 |
18 flagships scored · coverage: Grounding 15/18 · Tool-Use 14/18. Click a column to sort. Function-Calling (BFCL v4) is collected (2 so far) but withheld until the benchmark covers more of the frontier.
Inverted Vectara HHEM hallucination rate (100 − rate). HHEM scores faithfulness when summarizing a given document — a proxy for RAG grounding, not full open-domain factuality.
τ-bench multi-turn task completion with policy adherence (retail/airline). The hardest, sparsest metric — many models have no published score yet.
The two axes come from different evaluations with different methodologies and snapshot dates — read them as directional signals per axis, not a single combined ranking. Hover any score for its source.
* = value from an alternative benchmark than the column's canonical source (hover for which) — e.g. Command A+ has no Vectara HHEM entry, so its Grounding uses AA-Omniscience Non-Hallucination instead.
Origin badges (US · CN · EU · CA) make the geopolitics explicit: scored here are 11 US, 5 China, 1 EU, 1 Canada. The frontier is a US–China race at the top, while the enterprise/RAG axis is exactly where the EU/Canada players (Mistral, Cohere) compete — when the canonical boards list them.