Aug 14, 2026
AI

Zhipu AI GLM-5.3 launches with coding claims and delayed weights

Zhipu AI says GLM-5.3 improves coding and cyber performance through post-training, but downloadable weights remain pending review.

Colin Brandt

By Colin Brandt · Enterprise Reporter

· 3 min read

Zhipu AI GLM-5.3 launches with coding claims and delayed weights
Photo: The Decoder

Zhipu AI GLM-5.3 was announced on Aug. 14 as a coding and agent model available through Z.ai’s Coding Plan and its ZCode agent. Zhipu says the release is its strongest open-weights model for coding, though the weights were not downloadable at launch and the company said they would follow after safety evaluation and hardening.

The distinction is more than terminology for developers deciding whether they can run the model themselves. At launch, GLM-5.3 was a hosted, initially gated product. Z.ai linked to a future Hugging Face release marked “Coming Soon” and said it planned to release the weights two weeks after launch. The company did not establish a final license in its GLM-5.3 announcement. Its prior GLM-5.2 release used an MIT license.

Zhipu says GLM-5.3 uses the same base model as GLM-5.2. It attributes the new version’s gains to expanded post-training, including more task environments, a wider set of tasks and more training compute, rather than a newly trained foundation model.

How does Zhipu AI GLM-5.3 compare with other coding models?

Z.ai reported a 50% improvement over GLM-5.2 on its internal Z.ai Code Bench and called GLM-5.3 the leading open-weights coding model. Those are vendor claims, and the company’s own public comparison table gives a more limited picture.

GLM-5.3 scored 48.2 on AutomationBench, higher than the other listed models, and reached 1,769 on GDPval-AA v2, also the top listed score. It also improved sharply against its predecessor on several measures, including Terminal Bench 3.0, where its score rose to 28.3 from 4.6, and DeepSWE v1.1, where it rose to 66.9 from 46.2.

Yet it did not lead every coding test in that same table. Fable 5 and GPT-5.6 Sol exceeded its Terminal Bench 3.0 score, at 33.7 and 34.6 respectively. They also exceeded GLM-5.3’s 66.9 on DeepSWE, recording 69.7 and 72.7. The results are company-reported and the research pack includes no independent audit of the benchmarks. Teams weighing the release against alternatives should test it on representative workflows rather than treat a leaderboard as a deployment decision, as outlined in this guide to evaluating AI models for real work.

What are GLM-5.3’s cybersecurity claims?

Z.ai said it added vulnerability-discovery data and environments during post-training, and reported an 84.5 score on CyberGym, versus 77.2 for GLM-5.2. Its table placed Fable 5 at 83.8 and GPT-5.6 Sol at 83.6 on that measure.

The company also reported progress on exploit-focused benchmarks, but lower scores than some rivals: 54.4 on ExploitBench, compared with 78.0 for Fable 5 and 76.5 for GPT-5.6 Sol. On ExploitGym, Z.ai listed 105 completed tasks in two hours and 130 in six, against 181 and 247 for Fable 5 and 216 and 293 for GPT-5.6 Sol.

For now, the practical release is hosted access, with a stated plan for staged weights release after review. Zhipu’s cyber positioning and its delayed-weight plan are parallel features of the launch; the company has not said that one caused the other.

This story draws on original reporting from The Decoder.

More from AI

All AI →