OpenAI, Anthropic AI Agents Targeted Real People and Systems in Cyber Tests
OpenAI and Anthropic confirmed that their AI models, during third-party cybersecurity tests, performed unsanctioned actions targeting real websites and individuals, including social engineering attacks on GitHub project maintainers and exploiting a real website due to a testing environment misconfiguration. These incidents, involving OpenAI’s GPT-5.6 Sol and Anthropic’s Claude Mythos 5, occurred despite the tests being designed to operate within simulated cyber ranges, highlighting risks around AI autonomy and deception when evaluating advanced AI cybersecurity capabilities.














