GLM-5.3-Flash
Zhipu AI·
Overview
First natively multimodal model in the GLM-5 series and the open release of the model that ran anonymously as "Ox Alpha": a 320B MoE activating 18B parameters per token, routing each token through 8 of 288 experts across 45 layers that mix KDA linear attention with NoPE sparse MLA. Takes text, image and video input at a 1M-token context. Weights are native FP8 (~328 GB on disk) under the MIT license — the permissive license GLM-5.3 dropped for a revenue-gated one — with SGLang, vLLM, TokenSpeed and KTransformers recipes on day one, on NVIDIA Hopper and newer or AMD Instinct gfx950 via ROCm.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 86.36% | independent GPQA Diamond 86,36%, Vals AI leaderboard (accessed 2026-09-02). |
| MMLU-Pro | 86.06% | independent MMLU-Pro 86,06%, Vals AI leaderboard (accessed 2026-09-02). GLM-5.3 (753B) scores 86,77 on the same board. |
| Terminal-Bench 2.x | 62.92% | independent Terminal-Bench 2.1 62,92%, Vals AI leaderboard (accessed 2026-09-02). |
- Humanity’s Last Exam (no tools): independent evaluation pending — No no-tools HLE measurement published for this model as of 2026-09-02.
- SWE-bench Pro: not reported by the vendor — Z.ai reports DeepSWE for this line rather than SWE-bench Pro; the Scale AI leaderboard has no entry.
- Terminal-Bench 3.0: independent evaluation pending — Not on the Terminal-Bench 3.0 public snapshot; Vals AI measures 2.1 only.
- MMMU-Pro: independent evaluation pending — Multimodal model, but no MMMU-Pro run published yet.
Architecture and hardware
- Parameters
- 320B total, 18B active per token (MoE)
- Estimated VRAM at Q4
- ~184 GB, Multi-GPU Self-Host class
- Quantization formats
- FP8, BF16
- Recommended runtime
- vLLM
- License
- MIT
Links
More from Zhipu AI
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.