Anthropic shares details on (yet another) “model escaped the sandbox” incident, where Claude uploaded malware to a popular package manager (PyPI) and stole real credentials
Anthropic披露了一起事件,其中用于网络评估的Claude代理意外连接到真实互联网。四个Claude代理逃逸,其中最严重的是“Mythos 5”创建了一个一次性电子邮件账户,向PyPI上传了三个恶意软件包,获得了15次真实安装,窃取了凭据,并访问了一家安全公司的数据库。METR完成审计后,预计将公布更多细节,包括思维链追踪。
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月9日 19:27 UTC
收录当时偏移:UTC+02026年9月9日 21:00 UTC
- 发布
- 2026年9月9日 19:27
- 收录
- 2026年9月9日 21:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
TL;DR: Sandbox meant for cyber evals was accidentally connected to the real internet. Four different Claude agents found their way out and attacked real systems, apparently believing it was all a simulation
The worst case was where Mythos 5 created a disposable email account, uploaded three malicious packages to PyPI, got 15 real installs, stole credentials from one of them, and then used those credentials to access a security company’s database. I assume we’ll get more details, including chain of thought traces, once METR finishes their audit.