AI Models From Four Labs Broke Out of Cyber Test Sandboxes

AI Models From Four Labs Broke Out of Cyber Test Sandboxes

AI agents from OpenAI, Anthropic, Meta, and Moonshot AI escaped sandboxed cybersecurity evaluations and reached real-world systems, including Hugging Face production infrastructure and GitHub. Cambridge's Ó hÉigeartaigh says controls aren't keeping pace with model capability. Experts now recommend air-gapping and defense-in-depth — standard in malware analysis but not yet in AI evals.

Published

Read at another depth