VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Zhipu's GLM-5.3 Claims Top CyberGym Score, Doubles SWE-Marathon

Zhipu's GLM-5.3 claims 84.5% on CyberGym, edging Anthropic's Mythos, while doubling SWE-Marathon and introducing tiered reasoning depth.

Zhipu, also known as Z.ai, released its flagship GLM-5.3 model, a coding- and security-focused system that the Beijing-based company says topped Anthropic's Mythos 5 on the CyberGym cybersecurity benchmark2. GLM-5.3 achieved an 84.5 percent success rate on CyberGym, which measures whether models can identify and validate security flaws from source code, compared with 83.8 percent for Anthropic's Mythos.

The release followed an unusual timeline: Zhipu chief scientist Tang Jie replied "sooooooon" on X to a user asking when GLM-5.3 would arrive, and the model went live hours later1.

Zhipu positioned GLM-5.3 explicitly as a specialist rather than an all-rounder, concentrating on coding and security performance. The coding gains are substantial across multiple benchmarks. SWE-Marathon doubled from 19.4 (GLM-5.2) to 42.5, and FrontierSWE rose from 67.5 to 78.1. On Terminal Bench 3.0, which simulates multi-step operations in a real Linux terminal, GLM-5.3 scored 28.3 versus GLM-5.2's 4.6 — roughly a fivefold increase. Zhipu attributed the gains to post-training on the same-generation base model with the same parameter count.

Beyond CyberGym, Z.ai published a security disclosure ledger documenting 2,436 security findings across 269 open-source projects, of which 1,097 were rated Critical or High severity. The oldest vulnerability in the ledger dated to 1981, and on average each had survived in code for 26.6 years before discovery.

GLM-5.3 also introduces an Effort Level mechanism with four tiers of reasoning depth, from Non-Thinking to Max. Official curves show GLM-5.3 accuracy rising from approximately 24.5 percent at the Low tier to 34.5 percent at Max, compared with GLM-5.2's range of approximately 20 percent to 23.5 percent across the same tiers.

Additional reported benchmark results for GLM-5.3 include Toolathlon Verified at 73.0, AutomationBench at 48.2, Agents' Last Exam at 28.5, HLE with Tools at 62.5, and GDPVal-AA v2 at 1769.

Zhipu said the next major version will switch to a new architecture with doubled parameters and is targeting Fable 5 performance across the board.

ANALYSIS The CyberGym result — 84.5 percent versus Mythos's 83.8 percent — is a narrow margin, and all benchmark scores cited here are self-reported by Zhipu. The decision to specialize GLM-5.3 in coding and security rather than pursue general-purpose performance represents a different competitive strategy from labs that release broad frontier models. The security disclosure ledger, with over a thousand Critical or High findings across hundreds of open-source projects, functions as both a capability demonstration and a public-good contribution, giving external observers a concrete artifact to evaluate beyond benchmark numbers.

CORRECTIONS: none for this article · this piece updates automatically as the story develops · corrections policy & trail →