
Grok 4.5 cuts inference token output to a quarter of Opus, reshaping agent economics
xAI's Grok 4.5, announced July 8, averages 15,954 output tokens per SWE Bench Pro task versus Opus 4.8's 67,020 — a 4.2x reduction. For agentic workloads that chain dozens of tool calls, token efficiency matters as much as benchmark scores. That gap fundamentally changes the unit cost of running these systems at scale, independent of raw accuracy gains.
Published