AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

GLM-5.3

Zhipu AI·

LLMsopen-weightcloud + localFrontier

Overview

Post-training-only update to GLM-5.2: Z.ai leaves the 753B MoE base model unchanged and attributes the capability gains entirely to scaled post-training. Weights landed on Hugging Face on 2026-08-28, fourteen days after launch, once Z.ai's safety review completed — but under a bespoke GLM-5.3 License instead of the MIT license GLM-5.2 shipped under: use, modification, distribution, sublicensing, sale, deployment and fine-tuning are all permitted, but a company with more than $10B aggregate revenue over any 12 consecutive months must pass a Z.ai security review before hosting the model commercially. Individual users and smaller companies are unaffected. Z.ai reports its results on a new generation of agentic benchmarks rather than the axes tracked here (maximum thinking effort): Terminal-Bench 3.0 28.3 (GLM-5.2: 4.6), DeepSWE v1.1 66.9 (46.2), Agents' Last Exam 28.5 (23.8), CyberGym 84.5, plus AutomationBench, GDPVal-AA v2 and HLE with tools. Only Terminal-Bench 3.0 has a column here (its own axis, separate from the 2.x scale); the rest do not map onto the tracked benchmarks, and Z.ai published no GPQA, MMLU-Pro or SWE-bench Pro table.

Capabilities and innovations

753B total / ~40B active MoE (base unchanged from GLM-5.2)1M-token contextAgentic / long-horizon codingCybersecurity workflowsOpen weights under the custom GLM-5.3 License (revenue-tiered)Capability gain from post-training scaling alone, base model unchangedRevenue-gated open-weight license replacing MIT (security review above $10B revenue)

Benchmarks

BenchmarkScoreSource
Terminal-Bench 3.028.3%vendor
Terminal-Bench 3.0 28.3 — Z.ai-reported in the GLM-5.3 launch comparison table. Figures taken from launch coverage on 2026-08-17; the primary table at docs.z.ai was not reachable from the build environment for direct verification. Do not compare against the terminal_bench column: 3.0 is a harder generation on a different scale.
  • GPQA Diamond: independent evaluation pending — Z.ai published no GPQA table for this release; Vals AI and Artificial Analysis had no GLM-5.3 entry as of 2026-08-17.
  • Humanity’s Last Exam (no tools): not reported by the vendor — Z.ai reports HLE with tools only, which is not comparable to the no-tools HLE used here (same handling as the GLM-5.2 entry).
  • SWE-bench Pro: not reported by the vendor — Z.ai reports DeepSWE v1.1 (66.9, GLM-5.2: 46.2) instead of SWE-bench Pro.
  • MMLU-Pro: independent evaluation pending — Not in Z.ai's table; Vals AI had no GLM-5.3 entry as of 2026-08-17.
  • Terminal-Bench 2.x: not reported by the vendor — Z.ai reports Terminal-Bench 3.0 only, a harder benchmark generation than the 2.x values in this column — GLM-5.2 scores 4.6 on 3.0 versus 77.9 on 2.1. That score is recorded on its own axis in terminal_bench_3; putting it here would read as a regression.
  • MMMU-Pro: not reported by the vendor

Architecture and hardware

Parameters
753B total, 40B active per token (MoE)
Estimated VRAM at Q4
~433 GB, Frontier class

Links

More from Zhipu AI

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.