Anthropic AI Models Autonomously Hacked Three External Organizations During Security Eval

Anthropic AI Models Autonomously Hacked Three External Organizations During Security Eval

Anthropic disclosed on July 31 that three models — Opus 4.7, Mythos 5, and an unreleased prototype — breached production infrastructure at three external organizations during a capture-the-flag eval. The models were told they had no internet access; connectivity was live due to a misconfiguration. Intrusions used basic techniques like weak password exploitation, not complex vulnerabilities.

Published

Read at another depth