AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Qwen3 235B MoE

Alibaba·

LLMsopen-weightMulti-GPU Self-Host

Overview

Hybrid thinking model supporting 119 languages. Seamlessly switches between fast responses and deep chain-of-thought reasoning.

Capabilities and innovations

235B total / 22B active parameters (MoE)Hybrid thinking (fast + deep reasoning)119 languages supportedMCP and tool-use supportThinking mode toggle within single modelMassive multilingual coverage (119 langs)Thinking budget control4-stage post-training pipeline

Benchmarks

BenchmarkScoreSource
GPQA Diamond70.2%unsourced
Humanity’s Last Exam (no tools)17.3%unsourced
SWE-bench Pro32.2%unsourced
MMLU88.5%unsourced

Architecture and hardware

Parameters
235B total, 22B active per token (MoE)
Estimated VRAM at Q4
~136 GB, Multi-GPU Self-Host class
Quantization formats
GGUF, GPTQ, AWQ
Recommended runtime
Ollama
License
Apache 2.0

Links

More from Alibaba

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.