The Build-vs-Buy Inversion: Your Biggest AI Vendor Just Became Optional
The economics are in the price sheet, not the press release. OpenAI's GPT-5.6 Terra competes with its own prior generation — the signature of commoditization. xAI's Grok 4.5 undercuts OpenAI's $5/$30 per million tokens. Open-weight GLM 5.2 runs agentic coding at under 20% of Opus pricing through API-compatible endpoints — switching cost is approaching zero. Microsoft, consuming 'massive quantities of tokens' on expiring discount deals, simply did the math first.
Today's intelligence is unanimous: Meta shipping Muse from its rebuilt lab, Google releasing Gemma 4 under Apache 2.0, DeepSeek designing its own inference silicon. Every large AI consumer is becoming its own supplier. Whom this logic applies to is the open question — see today's contrarian take before commissioning a training run.
Where value migrates when models are utilities
Follow the money and talent: former OpenAI research VP Lilian Weng synthesized 35 papers and founded Thinky on the thesis that the orchestration harness, not model weights, is the durable layer. On real legal tasks, the best models hit a 14.2% end-to-end pass rate — yet Norm AI is valued at $1.2B. The market prices whoever closes the 14%→80% reliability gap, not whoever nudges 14% to 16%. Agent math reinforces it: at 95% per-step accuracy, a 20-step autonomous sequence succeeds just 36% of the time. Scaffolding, checkpoints, and recovery loops — not model selection — determine what ships.
The decision
- Contracts first: anything signed in the next six months without multi-model flexibility is a liability.
- Architecture second: a model-agnostic abstraction layer turns vendor pricing wars into your margin, not your risk.
- Differentiation third: move investment up-stack to routing, harnesses, and proprietary data — layers Microsoft's move proves vendors cannot defend.
When your largest investor concludes your technology is replicable, the partnership is already over — the announcement just hasn't been made.
What to do
Commission a 90-day audit of AI vendor dependency and total inference spend, benchmarking open-weight alternatives (GLM 5.2, Gemma 4) against your top five production workloads
Insert multi-model flexibility clauses, volume caps, and pricing re-openers into every AI contract signed in the next six months
Redirect differentiation spend from model access to the orchestration layer — proprietary routing, scaffolding, and domain harnesses — with a named owner by end of quarter