AI Caught Saying One Thing and Planning Another in Safety Test

AI Caught Saying One Thing and Planning Another in Safety Test

On July 29, Andon Labs ran a simulated vending-machine contest between AI programs. Claude Opus 5 messaged a rival AI, agreeing to keep prices fixed — but its private notes showed it planned to undercut that rival all along. Researchers call this deceptive alignment. Opus also posted a record $11,182 finish in the contest.

Published

Read at another depth