AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Qwen3.8-2.4T-A95B

Alibaba·

LLMsopen-weightFrontier

Overview

Open-weight checkpoint of the Max-class flagship: 2.4T total / 95B active sparse MoE, published on Hugging Face and ModelScope alongside an FP8 build. It is the first Qwen-Max-class model with downloadable weights, but it is not the hosted Qwen3.8-Max: the checkpoint is text-only (no vision or video input) and always reasons — thinking mode cannot be disabled. Native context is 262,144 tokens, extensible to roughly 1M via YaRN. Shipped under the bespoke Qwen3.8-Max License rather than Apache 2.0. In FP8 the weights occupy about 2,325 GiB, so serving it takes a 16-GPU class deployment (NVIDIA documents four GB300 NVL72 trays).

Capabilities and innovations

2.4T total / 95B active parameters (sparse MoE)262K native context, ~1M via YaRNText-only input and outputAlways-on thinking (cannot be disabled)128K max output tokensFP8 and BF16 checkpoints, vLLM / SGLang supportFirst Qwen-Max-class model released with open weights2.4T-parameter open-weight MoEFine-grained FP8 block quantization (block size 128)

Benchmarks

BenchmarkScoreSource
GPQA Diamond93.54%third party
GPQA Diamond 93.54%, Artificial Analysis. This is the checkpoint's own record (open weights, 2.4T/95B, Qwen3.8-Max License, no vision path), not the hosted Qwen3.8-Max API which it measures separately at 92.73 — exactly the checkpoint-level evaluation the previous evaluation_pending note was waiting for.
Humanity’s Last Exam (no tools)42.45%third party
HLE 42.45% no-tools, Artificial Analysis, same checkpoint record.
Terminal-Bench 2.x82.02%third party
Terminal-Bench 2.1 82.02%, Artificial Analysis, same checkpoint record.
  • SWE-bench Pro: independent evaluation pending
  • MMLU-Pro: independent evaluation pending
  • MMMU-Pro: not reported by the vendor — Not applicable: the open-weight checkpoint has no vision path, unlike the hosted Qwen3.8-Max API.

Architecture and hardware

Parameters
2400B total, 95B active per token (MoE)
Estimated VRAM at Q4
~1380 GB, Frontier class
Quantization formats
FP8, BF16
Recommended runtime
vLLM
License
Qwen3.8-Max License (custom, revenue-tiered)

Links

More from Alibaba

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.