The Routing Layer Just Got Its Price Discovery Event
Consensus called model routers thin wrappers with no defensibility; two independent cost proofs and one multibillion-dollar bid say the decision right is the asset.
What a router actually owns
A router owns no model. It owns the model-selection decision, which is where the spread between token volume and token spend gets captured, and along the way it accumulates something no single lab can replicate: a per-model, per-task performance table. Devshot's reporting puts numbers on how valuable that table is. Matching edit format to model swings agent success from 66% (DeepSeek, unified diff) to 94% (Doubao, JSON Patch). Reviewer pairing is asymmetric in the same direction. Claude checking Codex lifts pass rates from 71.6% to 89.7%, while Codex checking Claude drops accuracy from 91.4% to 82.8%. Those are not tuning footnotes. That is the routing table, and routing tables compound with usage.
Two independent cost proofs, one bid
The substitution here is measured rather than theoretical, which in this market is rarer than it sounds. Cursor's Auto Intelligence mode produces output users judge as good as the premium model at roughly 60% lower cost, per Exponential View's read of Vercel gateway data. Cursor separately cut a browser-building experiment from $10,000 to $1,300, about 87%, by routing routine work to cheap execution models behind a frontier planner. And Lenny's Newsletter reports a blind seven-model practitioner benchmark in which Claude Sonnet 5 scored 77 against Opus 5's 78, which is a vendor's own mid-tier model within one point of its flagship.
Then the price discovery. The Information reports OpenRouter fielding multibillion-dollar takeover interest. Strategic acquirers are paying for demand aggregation across models, cross-model performance data, and the switching friction that appears once a router sits in the critical path of production inference. The market spent two years dismissing that as a moat argument.
Where the sources genuinely disagree
The deflation story is not clean, and the disagreement is the useful part. Anthropic shipped Opus 5 at unchanged pricing of $5 per million input and $25 per million output tokens while claiming efficiency gains, passed none of them through as a price cut, and opened a paid latency tier at roughly 2.5x speed for 2x the rate. That is price discrimination, not a price war. ChinAI reports Moonshot going the other way entirely: Kimi K3 launched at $2.30 per million blended tokens, a 3.5x output-price increase over its own predecessor, thirteen times DeepSeek V4 Pro's $0.18.
If frontier prices stop falling, the automatic COGS tailwind penciled into your AI application models disappears — and the only remaining path to software-grade gross margin is architectural.
Two readings survive, and this is probably wrong, but the first is the one to underwrite. Either capability parity is converting into pricing power at the model layer, in which case app-layer margin expansion has to be engineered rather than waited for. Or Moonshot's increase is a capacity constraint wearing a strategy costume, since it suspended subscriptions within 48 hours of launch because demand exceeded compute. Allocation looks identical under both. Durable margin sits with whoever decides which model runs, not whoever trained it.
The reflex to avoid
The cheap version of this trade funds another gateway. The expensive version keeps crediting 'we use the best model' as defensibility while the routing question goes unasked in diligence. Founders who have already tested output parity at 60% traffic diversion are structurally cheaper to scale, and the re-architecture money not spent on them later is money available for the next position. Founders who have not have just told you something about their rigor.
What to do
Commission a routing-adjusted gross-margin case on every inference-exposed position this month, modelling a premium collapse from 8x to 3x revenue per token and 50%+ of volume diverted to cheap execution models.
Add one gating diligence question to every AI application memo: what is gross margin if 60% of traffic routes to non-frontier models, and has output parity been tested with human scoring?
Map five routing, model-selection and eval-observability teams for first meetings this quarter, prioritising those holding proprietary per-model performance data over integration surface.