GLM-5.3
Zhipu AI·
Overview
Post-training-only update to GLM-5.2: Z.ai leaves the 753B MoE base model unchanged and attributes the capability gains entirely to scaled post-training. Weights landed on Hugging Face on 2026-08-28, fourteen days after launch, once Z.ai's safety review completed — but under a bespoke GLM-5.3 License instead of the MIT license GLM-5.2 shipped under: use, modification, distribution, sublicensing, sale, deployment and fine-tuning are all permitted, but a company with more than $10B aggregate revenue over any 12 consecutive months must pass a Z.ai security review before hosting the model commercially. Individual users and smaller companies are unaffected. Z.ai reports its results on a new generation of agentic benchmarks rather than the axes tracked here (maximum thinking effort): Terminal-Bench 3.0 28.3 (GLM-5.2: 4.6), DeepSWE v1.1 66.9 (46.2), Agents' Last Exam 28.5 (23.8), CyberGym 84.5, plus AutomationBench, GDPVal-AA v2 and HLE with tools. Only Terminal-Bench 3.0 has a column here (its own axis, separate from the 2.x scale); the rest do not map onto the tracked benchmarks, and Z.ai published no GPQA, MMLU-Pro or SWE-bench Pro table.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| Terminal-Bench 3.0 | 28.3% | vendor Terminal-Bench 3.0 28.3 — Z.ai-reported in the GLM-5.3 launch comparison table. Figures taken from launch coverage on 2026-08-17; the primary table at docs.z.ai was not reachable from the build environment for direct verification. Do not compare against the terminal_bench column: 3.0 is a harder generation on a different scale. |
- GPQA Diamond: independent evaluation pending — Z.ai published no GPQA table for this release; Vals AI and Artificial Analysis had no GLM-5.3 entry as of 2026-08-17.
- Humanity’s Last Exam (no tools): not reported by the vendor — Z.ai reports HLE with tools only, which is not comparable to the no-tools HLE used here (same handling as the GLM-5.2 entry).
- SWE-bench Pro: not reported by the vendor — Z.ai reports DeepSWE v1.1 (66.9, GLM-5.2: 46.2) instead of SWE-bench Pro.
- MMLU-Pro: independent evaluation pending — Not in Z.ai's table; Vals AI had no GLM-5.3 entry as of 2026-08-17.
- Terminal-Bench 2.x: not reported by the vendor — Z.ai reports Terminal-Bench 3.0 only, a harder benchmark generation than the 2.x values in this column — GLM-5.2 scores 4.6 on 3.0 versus 77.9 on 2.1. That score is recorded on its own axis in terminal_bench_3; putting it here would read as a regression.
- MMMU-Pro: not reported by the vendor
Architecture and hardware
- Parameters
- 753B total, 40B active per token (MoE)
- Estimated VRAM at Q4
- ~433 GB, Frontier class
Links
More from Zhipu AI
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.