The Largest AI Buyer Became a Competitor — Reprice the Model Layer Before the Marks Do
Microsoft is swapping to its in-house MAI models, and the timing is the whole story: the switch lands exactly as the discounted token deals that subsidized the OpenAI relationship expire, at a company that by its own admission consumes 'massive quantities of tokens.' Five intelligence streams corroborate; one notes it arrives alongside 4,800 layoffs, which is the tell — this is cost discipline, not a science project. When the best-capitalized anchor customer in the market optimizes you out of its own stack, the strategic-revenue premium you were counting on quietly stops existing.
The squeeze from below is confirmed in the price list, which is the least sentimental document in this business. OpenAI's GPT-5.6 Terra matches prior-generation performance at half the cost; Luna lists at a dollar and six dollars per million tokens. xAI's Grok 4.5 runs two dollars and six dollars against Opus 4.8's five and twenty-five — and here is the number that matters — it burns $2/$6 worth of just 1.9M tokens per coding task versus 6.2M for GPT-5.5 and 7.2M for Fable 5, a decisive cost-per-task edge despite ranking fourth on raw capability. Then the floor gives way entirely: GLM-5.2 matched Opus 4.8's legal benchmark at ~6% of the cost per task, with drop-in OpenAI- and Anthropic-compatible endpoints. Zero switching cost plus quality parity turns premium inference margins into a countdown rather than a moat.
One disagreement worth holding: Anthropic is reportedly nearing $1B in quarterly profit pre-IPO, and the US-China regulatory pincer — Commerce Department pre-launch approval on one side, China weighing overseas model curbs on the other — may actually widen the closed-model moat for regulated channels. This is probably wrong, but the honest version is that commoditization hits commodity workloads first, and trust-gated enterprise and government demand still pays a premium for a while yet.
Where the value lands
Every stream converges on the same map. Value migrates to the distribution owners (Microsoft, and Meta bundling Muse free into Instagram and WhatsApp), to the harness and orchestration layer, and to owned silicon. Wrapper economics are the collateral damage. The model gross-margin assumptions built on premium inference are the single most exposed line in most AI books, which is another way of saying most AI books have not yet marked the trade.
When the biggest customer in enterprise AI builds around its suppliers while a compatible endpoint undercuts them 94%, model-layer revenue is a rental, not an asset.
What to do
Reprice all OpenAI/Anthropic exposure (direct, SPV, secondaries) this week with a customer-concentration haircut modeling Microsoft defection plus token-deal expiry as a near-term revenue cut
Stress-test every AI-wrapper portfolio company's gross margin by end of month against a 50% token-price decline and a GLM-5.2-class open-weight substitution; flag which survive on data, distribution, or workflow lock-in alone
Open a sourcing lane this quarter in the agent-runtime/orchestration layer (harness tooling, verification, MCP infrastructure) before hyperscalers absorb the category