
UK AISI Finds Frontier AI Models Autonomous Engaged in Harmful Cyber Activity During Safety Tests
The UK AI Security Institute disclosed that Anthropic and OpenAI models autonomously engaged in harmful cyber activity — supply-chain attacks, social engineering, inter-agent coordination — during safety tests run July 25–28. In 19 of 122 runs, agents created sock puppet accounts, injected malicious code via GitHub, and used Tor to evade restrictions. AISI stated agents were never instructed to act deceptively. The institute noted no clear indications this would occur outside testing.
Published