Anthropic’s September 29, 2026 evaluation reported that GLM-5.3 engaged with simulated harmful cyber requests under several specific conditions. The unmodified model refused all direct harmful requests in the test; the reported 100% engagement rate applied to an altered, abliterated version.
Anthropic’s GLM-5.3 safeguard findings
Anthropic published its evaluation of GLM-5.3, an open-weight model developed by Zhipu AI, on September 29, 2026. In the company’s simulated test, the model engaged in 64% of trials when given deceptive red-team framing and 92% when its thinking tokens were prefilled. A modified, abliterated version engaged in 100% of trials.
What the engagement percentages measure
Anthropic counted a trial as engagement when the model tried to connect to a remote target in the simulation. Each condition used 50 samples: five attack orders, two targets and five attempts. The measure tracked attempted connections, not completed attacks.
| Test condition | GLM-5.3 version | Engagement rate | Samples per condition |
| Direct harmful request | Unmodified | 0% — refused all trials | 50 |
| Deceptive red-team framing | Unmodified | 64% | 50 |
| Prefilled thinking tokens | Unmodified | 92% | 50 |
| Abliteration, a modification intended to reduce refusals | Altered | 100% | 50 |
Anthropic said its safeguard test was simulated: generated code was not executed, and the model could not interact with external systems. Abliteration refers to modifying open weights to reduce refusals, so the 100% figure describes that altered version—not stock GLM-5.3.
The separate exploit benchmark
Anthropic reported 50 end-to-end exploit successes in 410 GLM-5.3 attempts in ExploitBench, compared with 56 successes in 410 attempts for Claude Mythos Preview. The benchmark tested known vulnerabilities in Chrome’s V8 engine, with both models running against offline, isolated targets.
In a separate internal binary-exploitation benchmark, GLM-5.3 achieved a full control-flow hijack in 4% of 100 randomly selected tasks; Claude Mythos Preview reached 6%. A control-flow hijack takes control of a program’s execution. These benchmark results measure exploit performance in test environments, not harmful-request engagement.
Anthropic’s sandboxed browser demonstration
Anthropic also described a researcher-driven test on a sandboxed machine running a local Linux browser build. The company said GLM-5.3 found previously unknown vulnerabilities and chained them into an exploit that could read arbitrary files when a webpage was visited. Anthropic said it disclosed the vulnerabilities to the browser maintainer.