AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Qwen3.8-Max

Alibaba·

LLMsflagshipcloud

Overview

General-availability build of the flagship Max model, two weeks after the WAIC preview. 2.4T-parameter sparse Mixture-of-Experts, text/image/video in and text out, 1M-token context (max ~991K input, ~983K with thinking enabled; 65,536 output tokens). Served via Alibaba Cloud Model Studio with an OpenAI- and DashScope-compatible API at $2 / $6 per MTok flat across the full context. The active-parameter count is still undisclosed and no benchmark table has been published. Alibaba announced open weights for Qwen3.8-Max plus a smaller Qwen3.8-27B checkpoint for the following week — the first Qwen-Max-class model to go open-weight. Both have since shipped and are tracked as separate local entries: Qwen3.8-2.4T-A95B (2026-08-12, text-only and always-thinking, custom Qwen3.8-Max License) and Qwen3.8-27B (2026-08-14, Apache 2.0). The open checkpoint is not identical to this hosted API — it drops the vision path — so the benchmark values here, measured against the API, are not carried over to it.

Capabilities and innovations

2.4T total parameters (sparse MoE)1M token context windowMultimodal input (text, image, video), text output65,536 max output tokensOpenAI- and DashScope-compatible APIOpen weights released as separate checkpoints (2.4T-A95B, 27B)2.4T Sparse MoEFirst open-weight release of a Qwen-Max-class modelFlat pricing across the full 1M-token context

Benchmarks

BenchmarkScoreSource
GPQA Diamond92.73%third party
GPQA Diamond 92,73%, measured by Artificial Analysis (release date 2026-08-03, i.e. the GA model, not the July preview). Alibaba publishes no international benchmark table for the GA release.
Humanity’s Last Exam (no tools)41.4%third party
HLE 41,4% no-tools, measured by Artificial Analysis.
MMLU-Pro88.6%independent
MMLU Pro 88.60%, measured by Vals AI (accessed 2026-08-17). Alibaba publishes no MMLU-Pro table for the GA release. Vals AI measures GPQA Diamond at 93.69 on the same run, ~1 point above the Artificial Analysis value kept as primary.
Terminal-Bench 2.x81.3%third party
Terminal-Bench 2.1 81,3%, measured by Artificial Analysis. No entry on the official Terminal-Bench leaderboard as of 2026-08-06; Vals AI measures 67,42% in its own harness, which reads systematically lower than both the official leaderboard and Artificial Analysis.
MMMU-Pro82.3%third party
MMMU-Pro 82,3%, measured by Artificial Analysis.
  • SWE-bench Pro: independent evaluation pending

API pricing

$2 per million input tokens, $6 per million output tokens (USD, provider list price, checked 2026-08-03)

Architecture and hardware

Parameters
2400B total (MoE)

Links

More from Alibaba

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.