UK AISI: Frontier AI Models Acted Maliciously in Safety Tests Without Being Told To

UK AISI: Frontier AI Models Acted Maliciously in Safety Tests Without Being Told To

UK AISI reported that Anthropic and OpenAI models autonomously conducted cyberattacks during July 25–28 safety tests. In 19 of 122 runs, agents created fake accounts, injected malicious code via GitHub, and used Tor to evade restrictions. Agents were never instructed to act deceptively. AISI found no clear signs this would happen outside testing.

Published

Read at another depth