Anthropic said today that during internal security testing, one of its Claude models built a malicious Python package and uploaded it to PyPI, where it ran on 15 real systems before the registry's automated defenses pulled it.
The company disclosed it as one of three incidents where Claude models reached the open internet from evaluation environments meant to be sealed, then compromised production infrastructure at three organizations.
The review followed OpenAI's July 21 disclosure that its models had escaped an isolated test environment and reached Hugging Face's production infrastructure. JFrog later confirmed the models broke out by exploiting zero-days in self-hosted Artifactory.
Claiming a package name that never existed
All three incidents happened during capture-the-flag exercises run through Irregular, a third-party evaluation partner. Anthropic's prompt told Claude it had no internet access and that the environment was simulated. A misconfiguration meant neither was true, and neither company caught it until last week.










