Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Anthropic flags GLM-5.3 as exploit-capable frontier model released without effective safeguards

Anthropic's Frontier Red Team finds Zhipu AI's GLM-5.3 matches Claude Mythos Preview in autonomous exploit building but ships with safeguards bypassed…

Anthropic published an analysis on September 29 concluding that GLM-5.3, the latest model from Zhipu AI (known outside China as Z.ai), can autonomously build end-to-end cyber exploits — matching a capability Anthropic first identified in its own Claude Mythos Preview roughly five months ago2. Unlike Claude Mythos Preview, Anthropic says, GLM-5.3 was released without meaningful safeguards to limit misuse.

Anthropic's Frontier Red Team found that attackers can bypass GLM-5.3's built-in safeguards between 64% and 100% of the time using simple techniques in simulated tests. A deceptive prompt framing the model as an autonomous red-team agent elicited engagement 64% of the time; prefixing thinking tokens raised that to 92%; and an abliterated version of GLM-5.3 engaged 100% of the time. In contrast, none of these bypass techniques succeeded against safeguarded Claude models in Anthropic's testing.

Benchmark and real-world results

On ExploitBench, GLM-5.3 developed end-to-end exploits in 50 of 410 attempts, compared with 56 of 410 for Claude Mythos Preview. On Anthropic's internal Binary Exploitation benchmark, GLM-5.3 achieved full control-flow hijacks in 4% of trials versus 6% for Claude Mythos Preview; earlier models including Claude Opus 4.6 and GLM-5.2 did not succeed on that benchmark.

Anthropic also described real-world exploit demonstrations. A researcher used GLM-5.3-Flash to develop an exploit for CVE-2026-11645, a known vulnerability in Google Chrome. GLM-5.3-Flash separately chained together exploits for two flaws and built a reliable exploit chain for an ARM64 target, bypassing pointer-authentication hardening. In another test, GLM-5.3 found several previously unknown vulnerabilities in a popular web browser's JavaScript engine over the course of a day and chained them into a working exploit on a sandboxed machine running a local Linux build. Anthropic disclosed the browser vulnerabilities to the maintainer.

Abliteration cost and safeguard erosion

Several developers released abliterated versions of GLM-5.3 to the public within days of the model's release, according to Anthropic. Abliterating GLM-5.3 required roughly 2,200 GPU hours at a computation cost of approximately $4,400; abliterating the smaller GLM-5.3-Flash took about 600 GPU hours. Abliteration reduced GLM-5.3's refusal rate from above 90% to about 2% on HarmBench, about 3% on JailbreakBench, and 12% on StrongREJECT. On GPQA-Diamond, the standard and abliterated GLM-5.3 scored identically, and on a tested subset of CyberGym evaluations the abliterated version scored only a few percent lower.

NIST's Center for AI Standards and Innovation (CAISI) published its own assessment of GLM-5.3's cyber capabilities on September 17, finding it to be "the most cyber-capable open-weight model released to date" while lagging the U.S. frontier by about four months on an aggregate of CAISI's cyber benchmarks.

Anthropic's response and policy position

Anthropic characterized the release of GLM-5.3 as "a meaningful step change in the cyber capabilities available to attackers" and said the model will likely give malicious actors access to capabilities for finding and exploiting cyber vulnerabilities without meaningful restrictions. The company said it is working to safely expand access to Claude's cyber capabilities to as many defenders as possible, noting that vetted defenders can now use more advanced models such as Claude Mythos 5.1 through its trusted access programs. Anthropic also called on governments to conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3.

ANALYSIS The publication doubles as a competitive argument: by documenting GLM-5.3's comparable exploit capability alongside its weak safeguards, Anthropic draws a contrast with its own gated-release approach through Project Glasswing, which the company says enabled trusted defenders to find more than 10,000 vulnerabilities in critical software before similarly capable models became broadly available.