Z.ai Launches GLM-5.3 With Surprising Cyber Capabilities
Z.ai has launched GLM-5.3, an upgraded model that demonstrates unexpectedly powerful multi-stage cyber exploitation and coding capabilities achieved entirely through scaled-up post-training.

Z.ai has released GLM-5.3, an update to its flagship model that achieves significant performance gains purely through scaled-up post-training. Built on the same base model as GLM-5.2, the new version leverages the existing training stack, which includes the IndexShare long-context technique, the SAO reinforcement learning method, and the open-source slime framework. Z.ai scaled this setup by training the model on diverse, simulated professional work environments. The model is currently available via Z.ai's API and its GLM Coding Plan, with open-weights scheduled for release in late August 2026 following a two-week safety hardening period.
For developers, GLM-5.3 introduces a breaking change: the API now supports low, high, and max thinking effort levels, and thinking can no longer be disabled. This thinking capability translates to major benchmark improvements. On Terminal-Bench 3.0, the model jumped to 28.3 from GLM-5.2's 4.6. It reached 66.9 on DeepSWE v1.1 and 28.5 on the CLI variant of Agents' Last Exam. On the proprietary Z.ai Code Bench, GLM-5.3 achieved a 50% improvement over its predecessor, scoring 31.4% while using roughly 50,000 output tokens, outperforming Claude Opus 4.8's 29.5% at 120,000 tokens. However, it still trails Claude Fable 5, which scored 39.5%, as well as OpenAI's GPT-5.6 Sol on several public coding suites.
The most surprising evolution occurred in cybersecurity, where the model began autonomously planning multi-stage exploitation chains. On CyberGym, GLM-5.3 scored 84.5% compared to GLM-5.2's 77.2%. On ExploitBench, it scored 54.4%, more than doubling its predecessor's 24.4%. Under time-normalized budgets on ExploitGym, GLM-5.3 completed 105 tasks within two hours and 130 within six hours, compared to 29 and 39 for GLM-5.2, though it remains behind Mythos 5's scores of 181 and 247. In real-world testing, the model identified 2,436 vulnerabilities across 269 open-source projects, including 1,097 critical or high-severity bugs. These findings, which include a Linux kernel use-after-free, a WebKit memory-handling flaw, and a FreeBSD parameter-validation bug, are tracked on the Z.ai Security Disclosure Ledger, which currently holds 53 public CVEs and 2,383 embargoed entries.
This is our own summary of reporting by Unite.AI



