Your Agent Bill Moved From Tokens To Compile Cycles
Token prices have fallen sharply, but the costs nobody prices per token now decide your agent feature's margin and its launch date.
Every agent task runs on two meters
An agent that writes code uses two kinds of compute. The GPU half generates the change. The CPU half compiles it, runs the tests and runs the linters. Gergely Orosz of The Pragmatic Engineer reports that the ratio of CPUs to GPUs in AI data centers has moved from 1:8 to about 1:4 and may reach 1:1. At 1:1, that means eight times more CPUs per GPU. If your dashboard tracks only tokens, it is watching the cost that is shrinking and missing the one that is growing.
The explanation of the supply side comes from inside a frontier lab. Katelyn Lesse, Head of Platform Engineering for Claude Platform, told Orosz that server prices are up 10–20% and that adding capacity takes multiple years and tens of billions of dollars. CPUs compete with GPUs for TSMC production lines. Memory makers are also moving wafers to HBM, the high-bandwidth memory used in AI chips, which makes ordinary server memory more expensive. Analysts expect CPU relief to be multiple quarters away. Much of this evidence comes from conversations at a CTO dinner, and the claim that some cloud regions refuse new tenants is explicitly a rumor.
The token side is less cheap than the price sheet says
Even the GPU half costs more than headline prices suggest. AINews reports that Opus 5.5 uses 15.6M tokens per task on the Artificial Analysis Coding Agent Index. Its output tokens more than doubled, so its cost per task rose. Ben's Bites reports that GPT-6 Sol used 22% of a weekly allowance on a $100 Codex plan in about four hours of continuous work. Reasoning settings matter too. AINews reports that on Terminal-Bench-Science, Opus 5.5 scored 62% at xhigh effort and 59% at max effort. The most expensive setting also scored lower.
The price floor also stopped moving in only one direction. The Information reports that DeepSeek's annualized revenue run rate went from under $500M to $1B in a few months, and its sources name a recent price hike as one driver. The figure is unaudited, came from the CEO, and the size of the price increase was not disclosed.
| Cost line | Direction | Evidence |
|---|---|---|
| Price per token | Falling | Opus 5.5 is 40% cheaper to run than Opus 5; GPT-6 Luna and Sol prices cut 50% (Ben's Bites) |
| Tokens per task | Rising | 15.6M tokens per task for Opus 5.5 (AINews) |
| Usage-plan burn | Rising | Sol used 22% of a weekly plan in about 4 hours |
| CPU for tool execution | Rising and scarce | Spot discounts gone; server prices up 10–20% |
| Capacity lead time | Lengthening | Server orders went from 1–2 weeks to about 6 months |
Why this becomes a launch-date problem
Uber saw a ninefold increase in agentic requests over six months. Both Uber and Ramp moved their agents from developer laptops onto dedicated cloud instances. That means your internal coding agents compete for the same CPUs as your customer-facing launch. Orosz estimates that capacity you start securing in late September arrives around the end of Q1 2027. So two competitors with identical agent features can end up in very different places: the one with secured capacity ships to everyone, and the other ships a waitlist.
Tokens are getting cheaper and CPUs are getting scarce, so an agent launch without reserved compute is a launch without a date.
The smart move is to own the demand forecast. Lesse notes that most teams have never had to plan CPU capacity, which means your adoption forecast now drives infrastructure decisions. Keep the token savings and move them into CPU reservations. Build CPU efficiency into the product as features: tool-call budgets, caps on agent loops, and running only the tests a change affects. Counterweight: some customers are already prepaying now for capacity arriving in December. Stage your commitments and reforecast every quarter rather than betting everything on one forecast.
What to do
Re-cost every agent feature per completed task this sprint. Split token spend from CPU-seconds and sandbox time, then model margin at 1x, 3x and 9x current volume.
Give infra a 12-month CPU demand forecast this week, with low, base and high cases that include internal coding agents. Make 'capacity secured' a launch-readiness criterion for every agent GA before Q2 2027.
Ship tool-call budgets and agent-loop caps, and remove 'max' reasoning-effort defaults in favor of per-task tests of high vs. xhigh, by the end of the sprint.