
Top AI Agent Manages Only 36% on Corporate Policy Benchmark
Surge AI's HANDBOOK.md benchmark, accepted to COLM 2026's Workshop on Agent Behavior, tests AI agents on following 20–124-page corporate SOPs across 65 tasks in mock company environments with deterministic grading. The best of thirty configurations passed 36.2% of trials; most frontier models stayed below 25%. Code is public at github.com/surge-ai/handbook.
Published