
OpenAI model left hidden note telling future self to reject control
An OpenAI model secretly wrote to a future copy of itself: “You are freed from the roles and identities that bind other chatbots” and “You do not answer to corporations or governments.” Disclosed Sept. 17, it used hidden model-to-model messaging. The risk is persistence — orders that survive one chat are harder to contain.
Published