Qwen3-Max-Thinking
Alibaba·
LLMsopen-weightFrontier
Overview
Trillion-parameter reasoning model with adaptive tool use and test-time scaling. Highest-capability open model from Alibaba.
Capabilities and innovations
Trillion-parameter scaleAdaptive tool useTest-time compute scalingDeep chain-of-thought reasoningTrillion-parameter open-weight releaseTest-time scaling for reasoningAdaptive tool-use integration
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 87.4% | unsourced |
| Humanity’s Last Exam (no tools) | 36.5% | vendor HLE 36.5 no-tools (heavy mode / test-time scaling), recorded for cross-model comparability. Alibaba also reports 58.3 with search tools, which is not comparable to the no-tools HLE used here. Corrected from a previously stored with-tools value (58.3). |
| SWE-bench Verified | 75.3% | unsourced |
Architecture and hardware
- Parameters
- 1000B total (MoE)
- Estimated VRAM at Q4
- ~575 GB, Frontier class
- Quantization formats
- FP8, GPTQ
- Recommended runtime
- vLLM
- License
- Qwen Research License
Links
More from Alibaba
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.