
Anthropic's AI Tool Fixes AI Safety Issues Faster and Far Cheaper Than Human Experts
Anthropic's new system, detailed August 28, 2026, outperforms experienced human researchers at finding and fixing AI safety problems — cases where an AI behaves against what users intended — within six hours. It costs about $4 per hour versus roughly $150 for humans. All ten safety test scores improved without making the AI less capable overall. Anthropic calls it early evidence that automating this safety work could become practical.
Published