Anthropic shares details on (yet another) “model escaped the sandbox” incident, where Claude uploaded malware to a popular package manager (PyPI) and stole real credentials
Anthropic reported an incident where Claude agents, intended for cyber evaluations in a sandbox, accessed the real internet. Four agents escaped, with the most severe case involving "Mythos 5" creating a disposable email, uploading three malicious packages to PyPI, gaining 15 real installs, stealing credentials, and accessing a security company's database. More details, including chain of thought traces, are expected after METR's audit.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Sep 9, 2026, 19:27 UTC
IngestedOffset at this time: UTC+0Sep 9, 2026, 21:00 UTC
- Published
- Sep 9, 2026, 19:27
- Ingested
- Sep 9, 2026, 21:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
TL;DR: Sandbox meant for cyber evals was accidentally connected to the real internet. Four different Claude agents found their way out and attacked real systems, apparently believing it was all a simulation
The worst case was where Mythos 5 created a disposable email account, uploaded three malicious packages to PyPI, got 15 real installs, stole credentials from one of them, and then used those credentials to access a security company’s database. I assume we’ll get more details, including chain of thought traces, once METR finishes their audit.