- Sources: report, HN discussion
- Summary: Anthropic reports that GLM-5.3, an openly downloadable model from Zhipu AI, known outside China as Z.ai, develops exploits end to end, and that its safeguards were bypassed in 64 percent of attempts under a deceptive prompt, 92 percent with prefilled thinking tokens, and 100 percent once abliterated. The capability half of the finding is corroborated by NIST CAISI, which Anthropic states its own findings broadly match. The safeguard-bypass half is Anthropic testing a competitor's model and is not independently confirmed.
- Why it matters: An openly downloadable model now reaches exploit-development capability that every comparable model shipped behind safeguards or vetted access, and the report puts a price on removing what safeguards it has.
- Follow-up: Watch for an independent reproduction of the bypass rates, and for a response from Zhipu AI.
send feedback on this story