The Three-Layer AI Cost Crisis — Your Gross Margins Are About to Get Audited
The Cost Stack Just Became Three Landlords Deep
The enterprise AI cost thesis broke this week, and it broke from three directions at once. The mechanism matters, because the portfolio response differs by layer.
Layer 1: model provider costs are rising, not falling. GPT-5.5 ships a headline two-times price increase. The honest number, once completion-token effects are included, is 49-92% net cost inflation. OpenAI is explicitly pivoting from market share to margin. Anthropic killed flat-rate subscriptions the same week. GitHub is rebuilding its platform for 30x more load, which is what metered pricing produces at scale. A twenty-dollar Claude subscription burning hundreds to thousands in compute was never going to last. The subsidy ended.
Layer 2: the enterprise incumbents are imposing per-action tolls. ServiceNow's Action Fabric meters every action an external AI agent takes against its data. SAP now bans external agent access without endorsement. Workday and HubSpot are both moving to usage-based agent pricing. DataDog caps MCP requests at 5,000 daily. JPMorgan's Mark Murphy named it correctly: a tax on customers using outside AI agents.
Layer 3: the actual ROI does not cover the bill. KKR's Pete Stavros told Milken the AI earnings uplift across the portfolio is 5%, not 50%. That is the first honest number from a tier-one GP with every incentive to round up. Uber's CTO admitted they 'blew through' the AI budget after turning on agentic tools. Sierra's Bret Taylor called the ramp 'pricey before ROI.'
What This Means for Your Book
Every AI application company now faces a cost stack with two independent landlords who can raise rent unilaterally — the model provider and the data incumbent — selling into a buyer measuring single-digit ROI against double-digit cost increases. The math:
A company paying 49-92% more for inference, plus new per-action fees from ServiceNow or SAP, selling into a buyer measuring 5% uplift, is running negative gross margin on AI features that were supposed to defend ARR.
The a16z speedrun cohort confirms it at seed: founders are spending $300K/year on agents in lieu of engineers, with Cursor bills exceeding engineer salaries. Token spend and headcount are both rising at once. The 'AI replaces headcount' thesis is empirically dead at the company level while platform costs compound.
Where Sources Diverge
There is a reasonable counter-thesis, and it is worth stating cleanly before dismissing it. Inference costs fall fast enough that the flat-rate model looks retroactively sensible by late 2026. This is the $700B hyperscaler capex bet, that raw compute supply bends the curve. Several sources cite it as possible. None underwrite it as base case. The more honest framing, or rather the less flattering one, is that usage-based pricing sticks, wrappers pass costs through, and a meaningful cohort of AI-feature companies discover the feature that defended ARR is the feature that compressed the margin they were defending ARR to protect.
What to do
Run a pricing-model stress test across every AI-exposed portfolio company this week — model gross margin under metered pricing vs. current flat-rate, with 20-40% platform toll stack added
Mandate multi-model routing architecture for any portco with >30% API COGS from a single provider within 90 days
Re-underwrite every active LBO model to ≤5% AI EBITDA lift over 24 months; present adjusted returns at next IC
Source 3-5 MCP gateway / agent-governance startups for active diligence at pre-A to Series A