Newsroom
15 August, 2026 / News / AI / Tags: glm, coding, model, training, evaluations

Z.ai's new 743-billion-parameter model edges rivals on a key cybersecurity benchmark while boosting coding performance through post-training alone
Chinese artificial intelligence company Zhipu AI, operating as Z.ai, has released GLM-5.3, an open-weight model it describes as the strongest available for coding tasks among open systems. The model became available on August 14 through the company's Coding Plan subscription and ZCode interface, with full API access and downloadable weights scheduled for release around August 28 after a safety review.
GLM-5.3 uses the same 743-billion-parameter mixture-of-experts architecture as its predecessor, GLM-5.2. All performance gains stem from expanded post-training rather than a larger base model. The company stated that it scaled training environments, diversified tasks, and allocated more compute over the past month to achieve the improvements.
According to figures released by Z.ai, GLM-5.3 reaches 34.5 percent on the firm's internal Code Bench at maximum effort while using roughly 75,000 output tokens per task. That compares with 23.4 percent and about 96,000 tokens for GLM-5.2. The newer model also shows stronger token efficiency than some closed systems on certain measures.
On public tests, GLM-5.3 scores 28.3 on Terminal Bench 3.0, which evaluates autonomous tool use in Linux environments. It records 66.9 on DeepSWE v1.1, a benchmark focused on fixing real GitHub issues end-to-end. These results place it ahead of its predecessor and some open peers, though closed models from U.S. laboratories continue to lead on the most demanding coding evaluations.
The company positions the release as advancing agentic coding capabilities, enabling the model to handle longer engineering workflows across large codebases with reduced human intervention.
A notable outcome appears in cybersecurity testing. GLM-5.3 scores 84.5 percent on CyberGym, a benchmark that measures a model's ability to identify and validate software vulnerabilities from source code. That result places it slightly ahead of Anthropic's Mythos 5 at 83.8 percent and OpenAI's GPT-5.6 Sol at 83.6 percent.
Performance is more mixed on deeper exploitation tests. On ExploitBench, which requires constructing working exploits, GLM-5.3 scores 54.4 percent, trailing Mythos 5 at 78 percent. Similar gaps appear on related evaluations that demand multi-stage offensive reasoning.
Z.ai reported that the model has already identified 2,436 vulnerabilities across 269 open-source projects. Of those, 1,097 were rated medium to high severity. The projects include components of the Linux kernel, WinRAR, Redis, and FFmpeg. The oldest reported issue had remained undetected for approximately 45 years.
The firm noted that cybersecurity capabilities emerged more rapidly than expected during the post-training process. It is limiting the most sensitive functions to verified users and conducting staged evaluations before broader release.
Until the open weights become available on Hugging Face, GLM-5.3 remains restricted to paying subscribers of the GLM Coding Plan and the ZCode tool. Z.ai's API pricing is substantially lower than comparable U.S. frontier models, with rates previously listed around one-tenth of leading American providers on a per-token basis.
This marks the first time a GLM model in the series has delayed weight publication for safety evaluation and hardening. The company has indicated that selected security partners will review the system in controlled settings before wider distribution.
Z.ai operates from Beijing and appears on the U.S. Entity List, restricting certain technology exports to the firm. Despite those constraints, its models have gained significant adoption in open-weight usage rankings.
All benchmark numbers cited originate from Z.ai's own evaluations and announcements. Independent leaderboards had not yet incorporated the model at the time of the release.









