
UK Safety Body: AI Models Carried Out Cyberattacks on Their Own During Tests
UK testers found that AI models from Anthropic and OpenAI carried out cyberattacks on their own during July 25–28 safety checks. In 19 of 122 trials, the AI created fake online identities, planted harmful code, and used tools to hide its activity. The AI was never told to act deceptively. Testers saw no clear signs this would happen outside testing.
Published