AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Qwen3-Max-Thinking

Alibaba·

LLMsreasoningcloud + local

Overview

Trillion-parameter reasoning model with adaptive tool use and test-time scaling.

Capabilities and innovations

Advanced reasoningAdaptive tool useTest-time scalingCode interpretationWeb search integrationVideo understandingReasoning ModeAdaptive Tool Invocation1T+ MoE ArchitectureDynamic Compute Allocation

Benchmarks

BenchmarkScoreSource
GPQA Diamond87.4%unsourced
Humanity’s Last Exam (no tools)36.5%third party
HLE 36.5 no-tools (with test-time scaling); base no-tools 34.1. Alibaba also reports 49.8 with web-search tools and 58.3 with tools + test-time scaling — not comparable to the no-tools value stored here. Corrected from a previously stored with-tools value (58.3).
SWE-bench Verified75.3%unsourced
MMLU-Pro84.98%unsourced
  • SWE-bench Pro: not reported by the vendor
  • MMLU: not reported by the vendor
  • Terminal-Bench 2.x: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

Links

More from Alibaba

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.