AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Qwen3.8-2.4T-A95B

Alibaba·

LLMsopen-weightcloud + localFrontier

Overview

Open-weight checkpoint of the Max-class flagship and the first Qwen-Max-class model with downloadable weights: 2.4T total / 95B active sparse MoE, published on Hugging Face and ModelScope alongside an FP8 build and served via API. It is not the hosted Qwen3.8-Max — the checkpoint is text-only (no vision or video input) and always reasons, thinking mode cannot be disabled. Native context is 262,144 tokens, extensible to roughly 1M via YaRN. Shipped under the bespoke Qwen3.8-Max License rather than Apache 2.0. In FP8 the weights occupy about 2,325 GiB, so self-hosting takes a 16-GPU class deployment.

Capabilities and innovations

2.4T total / 95B active parameters (sparse MoE)262K native context, ~1M via YaRNText-only input and outputAlways-on thinking (cannot be disabled)128K max output tokensFP8 and BF16 checkpoints, vLLM / SGLang supportFirst Qwen-Max-class model released with open weights2.4T-parameter open-weight MoEFine-grained FP8 block quantization (block size 128)

Benchmarks

BenchmarkScoreSource
GPQA Diamond93.54%third party
GPQA Diamond 93.54%, measured by Artificial Analysis against the open-weight checkpoint itself (own record: open weights, 2.4T/95B, Qwen3.8-Max License), not against the hosted Qwen3.8-Max API which it measures separately at 92.73.
Humanity’s Last Exam (no tools)42.45%third party
HLE 42.45% no-tools, Artificial Analysis, same checkpoint record (hosted Qwen3.8-Max: 43.05).
Terminal-Bench 2.x82.02%third party
Terminal-Bench 2.1 82.02%, Artificial Analysis, same checkpoint record (hosted Qwen3.8-Max: 81.27). Same source and harness as the DeepSeek-V4 and Qwen3.8-Max entries.
  • SWE-bench Pro: independent evaluation pending — Alibaba published no benchmark table for the open-weight checkpoint, and neither Vals AI nor Artificial Analysis runs SWE-bench Pro.
  • MMLU-Pro: independent evaluation pending — Vals AI has no entry for this checkpoint as of 2026-08-17 (it does cover the hosted Qwen3.8-Max at 88.60).
  • MMMU-Pro: not reported by the vendor — Not applicable: the open-weight checkpoint has no vision path, unlike the hosted Qwen3.8-Max API. Artificial Analysis reports mmmuPro as null for it, which is consistent.
  • Terminal-Bench 3.0: not reported by the vendor

API pricing

$2 per million input tokens, $6 per million output tokens (USD, cheapest OpenRouter endpoint, checked 2026-08-17)

Architecture and hardware

Parameters
2400B total, 95B active per token (MoE)
Estimated VRAM at Q4
~1380 GB, Frontier class

Links

More from Alibaba

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.