AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

GLM-5.3-Flash

Zhipu AI·

LLMsopen-weightMulti-GPU Self-Host

Overview

First natively multimodal model in the GLM-5 series and the open release of the model that ran anonymously as "Ox Alpha": a 320B MoE activating 18B parameters per token, routing each token through 8 of 288 experts across 45 layers that mix KDA linear attention with NoPE sparse MLA. Takes text, image and video input at a 1M-token context. Weights are native FP8 (~328 GB on disk) under the MIT license — the permissive license GLM-5.3 dropped for a revenue-gated one — with SGLang, vLLM, TokenSpeed and KTransformers recipes on day one, on NVIDIA Hopper and newer or AMD Instinct gfx950 via ROCm.

Capabilities and innovations

320B total / 18B active MoE (8 of 288 experts)Native multimodal: text, image and video input1M-token contextHybrid KDA linear + sparse MLA attentionMIT licenseFirst natively multimodal GLM-5 modelHybrid KDA / sparse-MLA attention stackNative FP8 checkpointReturns the GLM line to MIT after the revenue-gated GLM-5.3 license

Benchmarks

BenchmarkScoreSource
GPQA Diamond86.36%independent
GPQA Diamond 86,36%, Vals AI leaderboard (accessed 2026-09-02).
MMLU-Pro86.06%independent
MMLU-Pro 86,06%, Vals AI leaderboard (accessed 2026-09-02). GLM-5.3 (753B) scores 86,77 on the same board.
Terminal-Bench 2.x62.92%independent
Terminal-Bench 2.1 62,92%, Vals AI leaderboard (accessed 2026-09-02).
  • Humanity’s Last Exam (no tools): independent evaluation pending — No no-tools HLE measurement published for this model as of 2026-09-02.
  • SWE-bench Pro: not reported by the vendor — Z.ai reports DeepSWE for this line rather than SWE-bench Pro; the Scale AI leaderboard has no entry.
  • Terminal-Bench 3.0: independent evaluation pending — Not on the Terminal-Bench 3.0 public snapshot; Vals AI measures 2.1 only.
  • MMMU-Pro: independent evaluation pending — Multimodal model, but no MMMU-Pro run published yet.

Architecture and hardware

Parameters
320B total, 18B active per token (MoE)
Estimated VRAM at Q4
~184 GB, Multi-GPU Self-Host class
Quantization formats
FP8, BF16
Recommended runtime
vLLM
License
MIT

Links

More from Zhipu AI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.