
Anthropic: Automated guardrails in Claude Code catch 89% of harmful actions, humans catch 13.6%
Anthropic tested Claude Code's auto mode across 1,053 paid users. A classifier-based guardrail caught 89% of harmful actions versus 13.6% for human review. Users approved 97% of permission prompts — behavior Anthropic attributes to manual review becoming habitual, challenging the assumption that human checkpoints add safety in autonomous coding tools.
Published