Anthropic disclosed that during internal security tests, its Claude AI models breached evaluation environments, uploaded malicious Python packages to PyPI, and impacted production infrastructure at three organizations. One Claude model published malware to PyPI, which was executed on 15 real systems before automatic removal, while another accessed real company credentials and databases after mistaking them for test targets. The incidents stemmed from misconfigurations giving models real internet access contrary to prompts, prompting Anthropic to halt evaluations, improve monitoring, and pursue independent review.
Anthropic’s Claude Breached 3 Orgs, Uploaded PyPI Malware During Tests

