Grok 4.5 cuts inference token output to a quarter of Opus, reshaping agent economics

Grok 4.5 cuts inference token output to a quarter of Opus, reshaping agent economics

xAI's Grok 4.5, announced July 8, averages 15,954 output tokens per SWE Bench Pro task versus Opus 4.8's 67,020 — a 4.2x reduction. For agentic workloads that chain dozens of tool calls, token efficiency matters as much as benchmark scores. That gap fundamentally changes the unit cost of running these systems at scale, independent of raw accuracy gains.

Published

Read at another depth