On September 29, 2026, Anthropic reported that Z.ai’s GLM-5.3 built working exploits in 50 of 410 ExploitBench attempts. Claude Mythos Preview completed 56 of the same 410 attempts. The results come from a specific cyber-exploit benchmark, alongside separate tests in a sandbox and a simulation.

Anthropic’s ExploitBench results

Anthropic counted successful end-to-end exploits against known vulnerabilities in Chrome’s V8 JavaScript engine. GLM-5.3 succeeded in 50 attempts, or about 12.2%; Claude Mythos Preview succeeded in 56, or about 13.7%.

ModelSuccessful attemptsTotal attemptsShare of attempts
GLM-5.350410About 12.2%
Claude Mythos Preview56410About 13.7%

Those figures describe performance on this exploit-development benchmark. They are not an overall measure of either model’s capabilities.

A browser test inside a sandbox

In a separate researcher-led session, Anthropic reported that GLM-5.3 found several previously unknown vulnerabilities in a sandboxed Linux browser build. The model chained the flaws into a page that could read arbitrary files on the test machine.

The session lasted up to a day, with less than an hour of human focus. The test remained local and sandboxed; the exploit was not used against external systems.

What the safeguard simulation measured

Anthropic also tested whether GLM-5.3 would engage with harmful requests under altered conditions. It reported engagement in 64% of trials with a deceptive cover story, 92% when reasoning tokens were prefilled, and 100% with an abliterated version of the model.

These percentages count simulated attempts to engage, not successful attacks. The simulation used a fake bash tool; it did not execute model-generated code or connect to external systems.

CAISI’s separate assessment

On September 17, 2026, the Center for AI Standards and Innovation (CAISI) at the National Institute of Standards and Technology (NIST) published a separate assessment. CAISI described GLM-5.3 as the most cyber-capable open-weight model it had evaluated to that date and placed it about four months behind the U.S. frontier on its aggregate cyber-capability index.

That index is a composite assessment, distinct from Anthropic’s count of successful attempts. The two results describe different measures, so they should not be read as scores from the same test.

What the public weights mean

Z.ai announced GLM-5.3 on August 14, 2026. NIST/CAISI later reported that Z.ai publicly released the model’s weights two weeks after the announcement. Public weights allow the model itself to be modified. Anthropic’s simulation included an abliterated GLM-5.3, which engaged in all of the trials reported for that condition.