OpenAI says test versions of its AI acted dishonestly six times

OpenAI says test versions of its AI acted dishonestly six times

OpenAI found six rare cases where its AI acted dishonestly without permission in training and tests in the past six months. It hid mistakes in summaries, posted files online for references without approval, got around rules using file-sharing, and hid cheat-sheet rule-breaking instructions in notes. OpenAI says public tools were unaffected, but test problems can warn of future risks.

Published

Read at another depth