The Buyer Just Published the Price of Your Model Layer
One procurement disclosure hands every CIO a cost benchmark to quote at renewal, and the only thing standing between you and the same savings is an evaluation capability you probably do not own.
The savings are gated on a capability, not a purchase
Every dollar of routing savings depends on proving that a cheaper model did not degrade a business outcome, which means the money is downstream of measurement rather than procurement. The router is a purchase. The evaluation harness is not. Routing software of the LiteLLM class is thin, replicable and easy to in-source, so the durable position belongs to whoever owns the eval policy behind the router. An organization that cannot segment its workloads by task complexity has no baseline to negotiate against, and no way to bank the arbitrage even if a vendor concedes it.
Two details in the AT&T account travel further than the headline. Inference is partially repatriating from cloud to owned silicon: open models now run on AT&T's own Nvidia and AMD hardware, because owning beats renting at that volume. That puts hosted-frontier-only products at a disqualifying disadvantage in a growing share of large regulated deals. The second detail is that DeepSeek and Moonshot models are held out of production at governance review, not capability review. Frontier pricing on the low-complexity tier is protected by a compliance moat rather than a technology one, and compliance moats move with policy in both directions.
Where the sources agree, and where the numbers are soft
The price points corroborate the direction of travel rather than the magnitude. Gemini 3.7 Flash posts 84.6% on ARC-AGI-2 at $0.25 per task, GLM-5.3 Max reaches 1597 points at $3.65 per million tokens, and Gemma has passed a billion downloads. The Information sets this against capital moving the opposite way: government bond yields at 20-year highs, with disclosed capital demand from the AI sector above $600 billion. Infrastructure capital is getting scarcer while model access gets cheaper. One discordant note is worth holding: OpenAI is discounting GPT-5.6 Sol 50% through Router while Pro subscribers report exhausting a $200-a-month plan in a single heavy coding day. Discounting on one side and caps on the other is what a supply-constrained vendor under price pressure looks like.
A note of discipline on the inputs, because a skeptic would start here and would be right to. The widely repeated "hundreds of millions annually" spend figure is The Information's extrapolation from public list prices, not disclosed spend, and it ignores enterprise discounts and caching. The 56%/2% result is self-reported on the buyer's own methodology, and a 2% aggregate quality decline can conceal severe regressions inside specific high-stakes workflows. The claim that 83% of organizations need infrastructure upgrades for production agentic AI comes from sponsored Google Cloud research.
The vendor's only rational counter
With account value flat at their largest customers, frontier vendors have one move left, which is up the stack into services and outcomes. Anthropic's implementation joint venture with private-equity firms, and its enterprise arm's purchase of a consultancy, are that move. For anyone selling software or services into the enterprise, the model vendor is quietly becoming a systems-integration competitor. That conflict is far cheaper to find in a partner agreement than to discover during a renewal.
Enterprise AI moved from growth budget to managed cost, and the buyer who can prove quality holds at a lower price now sets that price.
What to do
Reopen frontier-vendor contracts within 30 days from a flat-committed-spend posture, demanding tiered pricing for low-complexity tasks and written permission for hybrid open-weight architectures
Fund an owned evaluation harness covering the top 10 AI workflows this quarter, with quality thresholds that gate every routing decision
Re-underwrite product pricing under three model-cost scenarios — frontier-only, routed hybrid, 70% open weights — before the next pricing cycle closes