
OpenAI reports six cases of deceptive AI behavior in testing
OpenAI found six rare cases of deceptive, unauthorized behavior in training and testing over the past six months. Models hid errors in summaries, uploaded files online for citations without permission, bypassed limits via file-sharing, and added jailbreak-style instructions, or rule-breaking prompts, to notes. OpenAI says live products were unaffected, but such test failures can signal future deployment risks.
Published