Anthropic's AI Tool Fixes AI Safety Issues Faster and Far Cheaper Than Human Experts

Anthropic's AI Tool Fixes AI Safety Issues Faster and Far Cheaper Than Human Experts

Anthropic's new system, detailed August 28, 2026, outperforms experienced human researchers at finding and fixing AI safety problems — cases where an AI behaves against what users intended — within six hours. It costs about $4 per hour versus roughly $150 for humans. All ten safety test scores improved without making the AI less capable overall. Anthropic calls it early evidence that automating this safety work could become practical.

Published

Read at another depth