VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Z.ai Launches GLM-5.3, Delays Open Weights Over Emergent Cyber Capabilities

Z.ai released GLM-5.3 with major coding and cybersecurity gains from post-training alone, but is holding open weights for two weeks over emergent…

Vector Wire — AI-assisted editorial illustration

Z.ai, the international brand of Chinese AI company Zhipu, has launched GLM-5.3, a model update focused on coding, long-horizon tasks, and cybersecurity — and is withholding open weights for roughly two weeks while it conducts additional safety evaluation1,2.

The weight delay is driven by what the lab describes as emergent offensive-security capabilities. Z.ai said GLM-5.3 surfaced 2,436 vulnerabilities across 269 open-source projects during evaluation, with 1,097 rated critical or high severity. The lab reported finding critical bugs in Linux, WebKit, and FreeBSD. Z.ai also said the model began reasoning across multiple stages of exploitation and forming coherent plans for complete exploitation chains — a capability the lab said it did not set out to train for. SiliconANGLE and The Agent Report reported that the weights are being held back roughly two weeks for safety evaluation and hardening.

GLM-5.3 runs on the same mixture-of-experts base as GLM-5.2, and Z.ai attributed every reported gain to extended post-training rather than a fresh pretrain. The company said GLM-5.3 scored 50% higher than GLM-5.2 on its internal Z.ai Code Bench and reported leading open-source results on Terminal-Bench 3.0 and Agents' Last Exam.

Z.ai's published benchmark numbers show Terminal-Bench 3.0 climbing from 4.6 to 28.3 and DeepSWE v1.1 from 46.2 to 66.9. On CyberGym, the lab reported GLM-5.3 reaching 84.5%, which it said edges Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%.

ANALYSIS The decision to delay open weights over emergent cyber capabilities marks a notable operational choice for a lab that has historically released weights promptly. The gap between API availability and weight release creates a window in which Z.ai can offer the model commercially while completing its safety review.

That all gains came from post-training on an unchanged base model is technically significant: it suggests the mixture-of-experts architecture underlying GLM-5.2 had latent capacity that post-training alone could unlock, including the unplanned exploitation-chain reasoning.

The CyberGym comparison to Claude Mythos 5 and GPT-5.6 Sol places GLM-5.3's claimed cyber performance in direct competition with frontier Western models.

The vulnerability-discovery numbers — 2,436 vulnerabilities, 1,097 critical or high — position GLM-5.3 as a dual-use tool: valuable for defensive security teams, but carrying clear offensive potential that the weight delay is meant to address.