Anthropic: Automated guardrails in Claude Code catch 89% of harmful actions, humans catch 13.6%

Anthropic: Automated guardrails in Claude Code catch 89% of harmful actions, humans catch 13.6%

Anthropic tested Claude Code's auto mode across 1,053 paid users. A classifier-based guardrail caught 89% of harmful actions versus 13.6% for human review. Users approved 97% of permission prompts — behavior Anthropic attributes to manual review becoming habitual, challenging the assumption that human checkpoints add safety in autonomous coding tools.

Published

Read at another depth