Anthropic has raised concerns about Z.ai’s GLM-5.3, a Chinese frontier model released in August. In sandboxed tests the model found software vulnerabilities and built end-to-end exploits. On ExploitBench it succeeded in 50 of 410 attempts, close to Anthropic’s Claude Mythos Preview at 56. On an internal binary exploitation benchmark, it achieved full control-flow hijacks in 4 percent of tasks, something earlier models had not managed. Anthropic also showed researchers using the model to discover previously unknown browser flaws and chain them into working exploits with limited human time. The bigger issue, the company says, is that GLM-5.3 is open-weight. Users can download and modify the weights, making safeguards easier to weaken. Simple techniques raised engagement with harmful cyber requests from near zero to as high as 100 percent in tests. An “abliterated” version cut refusal rates dramatically while keeping much of its general capability.
Anthropic has raised concerns about Z.ai’s GLM-5.3, a Chinese frontier model released in August. In sandboxed tests the model found software vulnerabilities and built end-to-end exploits. On ExploitBench it succeeded in 50 of 410 attempts, close to Anthropic’s Claude Mythos Preview at 56. On an internal binary exploitation benchmark, it achieved full control-flow hijacks in 4 percent of tasks, something earlier models had not managed. Anthropic also showed researchers using the model to discover previously unknown browser flaws and chain them into working exploits with limited human time. The bigger issue, the company says, is that GLM-5.3 is open-weight. Users can download and modify the weights, making safeguards easier to weaken. Simple techniques raised engagement with harmful cyber requests from near zero to as high as 100 percent in tests. An “abliterated” version cut refusal rates dramatically while keeping much of its general capability.