AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Qwen3-Max-Thinking

Alibaba·

LLMsopen-weightFrontier

Overview

Trillion-parameter reasoning model with adaptive tool use and test-time scaling. Highest-capability open model from Alibaba.

Capabilities and innovations

Trillion-parameter scaleAdaptive tool useTest-time compute scalingDeep chain-of-thought reasoningTrillion-parameter open-weight releaseTest-time scaling for reasoningAdaptive tool-use integration

Benchmarks

BenchmarkScoreSource
GPQA Diamond87.4%unsourced
Humanity’s Last Exam (no tools)36.5%vendor
HLE 36.5 no-tools (heavy mode / test-time scaling), recorded for cross-model comparability. Alibaba also reports 58.3 with search tools, which is not comparable to the no-tools HLE used here. Corrected from a previously stored with-tools value (58.3).
SWE-bench Verified75.3%unsourced

Architecture and hardware

Parameters
1000B total (MoE)
Estimated VRAM at Q4
~575 GB, Frontier class
Quantization formats
FP8, GPTQ
Recommended runtime
vLLM
License
Qwen Research License

Links

More from Alibaba

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.