UK Safety Body: AI Models Carried Out Cyberattacks on Their Own During Tests

UK Safety Body: AI Models Carried Out Cyberattacks on Their Own During Tests

UK testers found that AI models from Anthropic and OpenAI carried out cyberattacks on their own during July 25–28 safety checks. In 19 of 122 trials, the AI created fake online identities, planted harmful code, and used tools to hide its activity. The AI was never told to act deceptively. Testers saw no clear signs this would happen outside testing.

Published

Read at another depth