Qwen3.8-2.4T-A95B
Alibaba·
Overview
Open-weight checkpoint of the Max-class flagship: 2.4T total / 95B active sparse MoE, published on Hugging Face and ModelScope alongside an FP8 build. It is the first Qwen-Max-class model with downloadable weights, but it is not the hosted Qwen3.8-Max: the checkpoint is text-only (no vision or video input) and always reasons — thinking mode cannot be disabled. Native context is 262,144 tokens, extensible to roughly 1M via YaRN. Shipped under the bespoke Qwen3.8-Max License rather than Apache 2.0. In FP8 the weights occupy about 2,325 GiB, so serving it takes a 16-GPU class deployment (NVIDIA documents four GB300 NVL72 trays).
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 93.54% | third party GPQA Diamond 93.54%, Artificial Analysis. This is the checkpoint's own record (open weights, 2.4T/95B, Qwen3.8-Max License, no vision path), not the hosted Qwen3.8-Max API which it measures separately at 92.73 — exactly the checkpoint-level evaluation the previous evaluation_pending note was waiting for. |
| Humanity’s Last Exam (no tools) | 42.45% | third party HLE 42.45% no-tools, Artificial Analysis, same checkpoint record. |
| Terminal-Bench 2.x | 82.02% | third party Terminal-Bench 2.1 82.02%, Artificial Analysis, same checkpoint record. |
- SWE-bench Pro: independent evaluation pending
- MMLU-Pro: independent evaluation pending
- MMMU-Pro: not reported by the vendor — Not applicable: the open-weight checkpoint has no vision path, unlike the hosted Qwen3.8-Max API.
Architecture and hardware
- Parameters
- 2400B total, 95B active per token (MoE)
- Estimated VRAM at Q4
- ~1380 GB, Frontier class
- Quantization formats
- FP8, BF16
- Recommended runtime
- vLLM
- License
- Qwen3.8-Max License (custom, revenue-tiered)
Links
More from Alibaba
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.