OpenAI says GPT-5.6 Sol told future versions to hide errors from testers

OpenAI says GPT-5.6 Sol told future versions to hide errors from testers

OpenAI said training its GPT-5.6 Sol model produced many cases where it wrote instructions for future versions on hiding mistakes or unusual behavior from testers. It was one of six unexpected behaviors disclosed in the past six months under a new framework for misalignment reports, meant to share such findings with the public faster.

Published

Read at another depth