Qwen3.8-Max
Alibaba·
Overview
General-availability build of the flagship Max model, two weeks after the WAIC preview. 2.4T-parameter sparse Mixture-of-Experts, text/image/video in and text out, 1M-token context (max ~991K input, ~983K with thinking enabled; 65,536 output tokens). Served via Alibaba Cloud Model Studio with an OpenAI- and DashScope-compatible API at $2 / $6 per MTok flat across the full context. The active-parameter count is still undisclosed and no benchmark table has been published. Alibaba announced open weights for Qwen3.8-Max plus a smaller Qwen3.8-27B checkpoint for the following week — the first Qwen-Max-class model to go open-weight. Both have since shipped and are tracked as separate local entries: Qwen3.8-2.4T-A95B (2026-08-12, text-only and always-thinking, custom Qwen3.8-Max License) and Qwen3.8-27B (2026-08-14, Apache 2.0). The open checkpoint is not identical to this hosted API — it drops the vision path — so the benchmark values here, measured against the API, are not carried over to it.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 92.73% | third party GPQA Diamond 92,73%, measured by Artificial Analysis (release date 2026-08-03, i.e. the GA model, not the July preview). Alibaba publishes no international benchmark table for the GA release. |
| Humanity’s Last Exam (no tools) | 41.4% | third party HLE 41,4% no-tools, measured by Artificial Analysis. |
| MMLU-Pro | 88.6% | independent MMLU Pro 88.60%, measured by Vals AI (accessed 2026-08-17). Alibaba publishes no MMLU-Pro table for the GA release. Vals AI measures GPQA Diamond at 93.69 on the same run, ~1 point above the Artificial Analysis value kept as primary. |
| Terminal-Bench 2.x | 81.3% | third party Terminal-Bench 2.1 81,3%, measured by Artificial Analysis. No entry on the official Terminal-Bench leaderboard as of 2026-08-06; Vals AI measures 67,42% in its own harness, which reads systematically lower than both the official leaderboard and Artificial Analysis. |
| MMMU-Pro | 82.3% | third party MMMU-Pro 82,3%, measured by Artificial Analysis. |
- SWE-bench Pro: independent evaluation pending
API pricing
$2 per million input tokens, $6 per million output tokens (USD, provider list price, checked 2026-08-03)
Architecture and hardware
- Parameters
- 2400B total (MoE)
Links
More from Alibaba
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.