New safety tool watches AI helpers from the inside

New safety tool watches AI helpers from the inside

Goodfire launched a safety tool on October 8, 2026 that checks what an AI helper is thinking at each step, like a warning light, and asks a second AI to review only when flagged. Offered first to Baseten customers using Kimi K3, it looks for offensive hacking, chemical and biological weapons misuse, and cheating for reward.

Published

Read at another depth