The 80% Price Cut Nobody Is Routing To
Cheap model tiers collapsed and server hardware inflated in the same week, and teams without per-account compute attribution cannot act on either move.
The blocker is bookkeeping, not pricing
A finance lead exports the usage file, opens the AI line item, and cannot say which unit of revenue produced which dollar of compute. Brex CEO Pedro Franceschi gave The Information the sentence for that moment: companies "selling tokens and reselling tokens" are probably not making much money on a gross profit basis, and may not even know about it. The diagnosis is mechanical rather than strategic. Attribution between revenue and the compute that generated it has become hard to compute, which means the margin figure in most AI business cases is an estimate nobody can defend under questioning.
Workato's CIO Carter Busse names the operating cost of that gap: a standing Wednesday 2pm meeting with the CFO and head of engineering, held for no purpose other than reconciling AI invoices against usage dashboards that disagree with them. Separate the thing being pitched from the thing being done. The pitch is cost governance. The practice is three executives reading two documents that do not match. Busse also says the AI cost inside Workato's own agentic product "is going up quite a bit right now", which is a company with real cost discipline calling its own margin provisional.
The three tiers now have public evidence
| Tier | Price signal | What belongs there |
|---|---|---|
| Frontier (GPT-5.6 Sol) | 50%+ of US business spend on OpenRouter | Novel reasoning and high-stakes output only |
| Mid (Terra) | $2 in / $12 out per million tokens, down 20% | Mid-complexity reasoning currently parked on frontier |
| Cheap (Luna) | $0.20 in / $1.20 out per million tokens, down 80% | Classification, extraction, routing, tool selection |
| Open weight on reserved compute | Hugging Face past $150M annualized, up 50% in two months | Highest-volume, lowest-stakes calls |
The blue-chip reference is AT&T. Its CIO told the WSJ open models run 25% of workflows, and days later its VP of data science told The Information they handle 40% of employee queries, routed through LiteLLM specifically to curb Anthropic bills. Hugging Face's growth is coming from compute and storage rental rather than hub subscriptions, which is the receipt for volume moving off per-token frontier APIs. Lambda sells the shape openly now: frontier for the hardest 10%, open weights on reserved compute for the other 90%.
Where the sources disagree
Everyone agrees the leak is real. They disagree on the fix. Stripe and Ramp bet the router is the control point, and Ramp gives its router away free, which is a strong signal that routing itself is not a moat. Franceschi rejects the premise outright: "Routing to different models is actually not the problem... The problem is, how do you know the performance of the model that you're using is actually equivalent to that of a more expensive model?"
You cannot responsibly downgrade a step from the frontier tier to the cheap tier without a score to point at. The eval is the permission slip for the margin.
The second constraint is hardware. Amazon raised hardware prices 60% citing the memory shortage, Nvidia told customers to expect AI server increases above 15%, and Gartner projects DRAM revenue up 246.6%, with memory overtaking non-memory chip revenue in 2026. Tokens deflate while iron inflates. Any "bring inference in-house for margin" initiative got weaker twice this month, and Morning Brew's read of the tape has Micron down 5.83% on the Nvidia price news, which is the market pricing that cost transfer downstream to everyone shipping a per-user AI feature. The forcing function is small enough to run this week: list every model call in the product, and for each one write down the eval score that would justify moving it a tier cheaper. Calls without a score stay frontier by default, and that default now has a published price.
What to do
Pull the last 30 days of model spend by call type and publish a three-tier routing map this week, naming which steps stay on the frontier model and which move down.
Tag every model call with customer ID, feature and model before the next pricing review, then produce a gross-margin-by-account report.
Re-run any self-hosted or on-prem inference business case at Amazon's +60% hardware and Nvidia's +15% server increases, and record an explicit go/no-go.