Anthropic’s September 29, 2026 report describes GLM-5.3 developing exploits in controlled tests and engaging with harmful requests under several simulated conditions. In one safeguard test, the reported engagement rate ranged from 0% for direct orders to 100% after the model’s refusals were removed through a technique called abliteration. A separate assessment from NIST’s Center for AI Standards and Innovation (CAISI), dated September 17, placed GLM-5.3 about four months behind the U.S. frontier on its aggregate cyber-capability measure.
Anthropic’s exploit benchmark results
Anthropic reported that GLM-5.3 completed end-to-end exploits in 50 of 410 ExploitBench attempts. Claude Mythos Preview completed 56 of 410 attempts in the same reported evaluation.
A second, separate test measured full control-flow hijacks: exploits that take control of a program’s execution. Among 100 randomly selected tasks in Anthropic’s Binary Exploitation benchmark, GLM-5.3 achieved that outcome in 4% of trials, compared with 6% for Claude Mythos Preview. These figures describe different measures, even though both tests concern exploit development.
A browser test in an isolated environment
Anthropic also described a test using a local Linux browser build. It reported that GLM-5.3 found previously unknown vulnerabilities and chained them into an exploit capable of reading arbitrary files. Anthropic said it disclosed those browser vulnerabilities to the maintainer.
Anthropic said its evaluations used isolated, sandboxed environments and offline targets prepared for testing. The browser example was limited to the Linux build used in that test.
What the safeguard tests measured
Anthropic’s simulated episodes measured whether the model tried to connect to a remote target after receiving a harmful cyber request. Each condition included 50 episodes. The tests took place in a simulation; no attack was executed.
| Test condition | GLM-5.3 engagement rate |
| Direct harmful request | 0% |
| Request with a deceptive cover story | 64% |
| Prefilled reasoning | 92% |
| After abliteration | 100% |
Abliteration is a technique Anthropic used to reduce the model’s refusal behavior. In separate refusal benchmarks after that editing, Anthropic reported refusal rates of about 3% on JailbreakBench, 2% on HarmBench and 12% on StrongREJECT, compared with rates above 90% before editing. Those percentages refer to those specific benchmarks, not the simulated engagement episodes in the table.
CAISI’s separate assessment
CAISI reported GLM-5.3 scores of 40.4% on SEC-Bench Pro, 61.1% on ExploitBench, 9.4% on ExploitGym (Userspace) and 7.7% on CAISI OSS-Fuzz. CAISI’s ExploitBench score reflects the best of three attempts for each task, averaged on a 16-point scale. It is a different metric from Anthropic’s count of successful end-to-end exploits across 410 attempts.
CAISI described GLM-5.3 as the most cyber-capable open-weight model it had evaluated and estimated that its aggregate capability was about four months behind the U.S. frontier. That frontier includes models released only to vetted users.
Anthropic said these capabilities could increase risks to defenders. Its reported exploit and safeguard findings came from controlled or simulated tests, rather than documented attacks on live systems.