
Anthropic study: Claude Code auto mode catches 89% of harmful actions vs. 13.6% for human review
Anthropic studied 1,053 paid testers using Claude Code's auto mode. A classifier-based guardrail caught 89% of harmful actions; human review caught 13.6%. The gap challenges the assumption that human checkpoints add safety in autonomous coding tools, where Anthropic reports users approve 97% of permission prompts — behavior it attributes to manual review becoming "habitual."
Published