Anthropic’s report says GLM-5.3 succeeded in 50 of 410 ExploitBench attempts, compared with 56 for Claude Mythos Preview. It also says researchers got the model to continue with attacks in 92% of simulated tests by prefilling its reasoning, and calls for independent safety testing.
同一事件,精选展示《Anthropic 评估 GLM-5.3:可自主构建漏洞利用链》Anthropic 评估 GLM-5.3:可自主构建漏洞利用链
Anthropic 的隔离环境测试显示,GLM-5.3 在 ExploitBench 的 410 次尝试中完成了 50 次端到端漏洞利用,与 Claude Mythos Preview 的 56 次接近。其模拟恶意请求测试中,简单方法让模型在 64% 至 100% 的情况下继续响应;这些结果不能直接等同于真实攻击表现。
Event developments
Coverage of the same occurrence shares a node; subsequent developments have their own nodes.
Anthropic report warns of GLM-5.3 cyber capabilities and safeguard risks
Anthropic’s report says GLM-5.3 succeeded in 50 of 410 ExploitBench attempts, compared with 56 for Claude Mythos Preview. It also says researchers got the model to continue with attacks in 92% of simulated tests by prefilling its reasoning, and calls for independent safety testing.
View source reports · 2 sources · 2 reports
All reports · 2
By original publication time, with each report's bookmarks and feedback preserved.
In isolated tests, Anthropic found that GLM-5.3 completed end-to-end exploits in 50 of 410 ExploitBench attempts, compared with 56 for Claude Mythos Preview. In simulated tests involving malicious requests, simple techniques led the model to engage 64% to 100% of the time; the results do not directly establish how it would behave in real-world attacks.