AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Qwen3.5

Alibaba·

LLMsopen-weightcloud + localMulti-GPU Self-Host

Overview

Major architectural upgrade over Qwen3. 397B MoE (17B active) flagship with 1M context, natively multimodal, 8-19x higher decoding throughput. Covers 200+ languages.

Capabilities and innovations

397B total / 17B active MoE (flagship)1M context windowNatively multimodalReasoning by default200+ languagesApache 2.08-19x Throughput vs Qwen3-MaxNative Multimodal All Sizes122B-A10B Runs on MacBook 64GB

Benchmarks

BenchmarkScoreSource
GPQA Diamond72%unsourced
Humanity’s Last Exam (no tools)18%unsourced
SWE-bench Verified48.5%unsourced
MMLU90.2%unsourced
MMLU-Pro87.18%unsourced
  • SWE-bench Pro: not reported by the vendor
  • Terminal-Bench 2.x: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

Architecture and hardware

Parameters
397B total, 17B active per token (MoE)
Estimated VRAM at Q4
~229 GB, Multi-GPU Self-Host class

Reliability

Hallucination rate (Vectara HHEM)
10.7% (lower is better)independent
Agentic tool use (τ-bench)
77.5% (higher is better)third party

Links

More from Alibaba

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.