Best AI Agent Passes Only 36% of Corporate Policy Benchmark Trials

Best AI Agent Passes Only 36% of Corporate Policy Benchmark Trials

Surge AI's HANDBOOK.md benchmark, accepted to the Workshop on Agent Behavior at COLM 2026, tests AI agents on following 20–124-page corporate SOPs across 65 tasks in mock company environments with deterministic grading. The best of thirty configurations passed 36.2% of trials; most frontier models stayed below 25%. Code is public at github.com/surge-ai/handbook.

Published

Read at another depth