Inference Deflated, Capacity Inflated — and Only One Shows Up in Your Renewal
Two AI cost curves are now moving in opposite directions, and any leader who nets them into a single line item will underprice the compute plan and overpay the model bill.
The two curves land on different lines of your P&L
Substitutable work is getting cheaper fast. Uber has held total AI cost flat since March while token and session volume kept climbing, with cost per request down 34% and cost per session down 52%, per The Pragmatic Engineer's account of the company's disclosures. On Uber's own numbers, its most expensive open-weight model runs $0.30 per code review against $0.50 for the cheapest frontier model and $2.50 for the most expensive. SpaceX reportedly cut Claude Code tokens by 90%. None of that required new science; it required a gateway between the application and the model plus a weekly benchmark.
Scarce capacity is moving the other way. Oracle's quarter, as reported by The Information, is a financing event rather than a demand event: management confirmed that most of the additional $30 billion of AI compute deals booked in the period came through customer prepayment or bring-your-own-chip structures. The market has quietly reorganized so that customers capitalize their vendor's build-out.
Capacity procurement is now a treasury decision
| Structure | Cash timing | Unit cost trajectory | Counterparty exposure |
|---|---|---|---|
| Traditional rental | Pay as consumed | Worst — repricing at +20% on renewal | Low; walk away at term |
| Prepaid capacity | Large upfront outlay | Best — you lock today's price | High; your cash sits on their balance sheet |
| BYO-chip colocation | Capex now, low opex later | Insulated from rental premiums | Medium; hardware is yours, siting is not |
Note the trap in row two. Prepayment is the only structure that reliably caps unit cost, and it maximizes exposure to a vendor S&P downgraded in July over OpenAI concentration and capex, whose stock is down 53% over twelve months and which is actively managing data center delays. Magouyrk's line that "all our eggs are not in a single basket" is site-diversification messaging, not a delivery guarantee.
What the 75x spread does not include
The log10.io figures are verified list prices per run and exclude self-hosting, GPU capacity, ops headcount, model validation, and regulated qualification work. The loaded multiple for a self-hosted open-weight deployment lands well below 75x and is still large enough to reorder a budget — but nobody publishing a number knows where. Two further constraints travel with the benchmark: the 0.6-point gap between first and second place is inside judge-panel noise, and the two labs driving the open-weight surge are Chinese-origin, with export-control and customer-contract review nowhere in the analysis. The enterprise savings figures, separately, are self-reported by the companies claiming them.
Where the evidence agrees, and where it diverges
Every credible account reviewed converges on the same architectural answer: a provider-agnostic routing layer, bought rather than built for version one, with an explicit quality-regression budget — AT&T's measured 2% is a defensible ceiling. The divergence is the interesting part. Ramp's telemetry shows AI spend falling 10% in August at the largest buyers, while Oracle books record AI bookings and OpenAI refuses revenue for lack of capacity. Both are true, and together they describe the actual market: buyers are driving down what they pay per unit of routine work while paying more for the scarce top tier.
The frontier premium is now a negotiable line item, and the only part of the compute bill that is still rising is the part you cannot substitute.
What to do
Rebuild the three-year compute forecast at flat-to-plus-20% unit pricing and publish a per-SKU AI contribution margin table with breakeven pricing before the next board cycle.
Reopen frontier model contracts ahead of renewal using the published AT&T and Pinterest figures as pricing anchors, and refuse any commitment past twelve months without a benchmarking or price-adjustment clause.
Commission one fully loaded self-hosted open-weight TCO — capacity, validation and qualification included — alongside an export-control and customer-contract review of Chinese-origin weights, this quarter.