The Margin Tailwind Went Free and the Input Went Scarce
Every agent company underwritten in the last 18 months assumed permanent token deflation, and both halves of that assumption have broken in opposite directions.
The term that matters is revenue share, not price
A per-token price increase is a cost problem, and cost problems have engineering answers: quantize the model, cache aggressively, shrink the tool surface, route the easy turns to something cheaper. A revenue-share demand is a different instrument entirely. Alibaba is reportedly weighing exactly that ask, per Benedict Evans, which converts an application company's gross margin from a number its engineers control into a number its supplier reopens at will. Anyone who underwrote a 70% steady-state margin behind a vendor holding revenue-share optionality underwrote a counterparty's forbearance and filed it as a cost curve.
Evans reads the market as a supply crunch in which labs can name their price. The counterweight he reports from the buy side is that large enterprises now presume dual-sourcing, because, in his phrase, you cannot rely on anyone right now. Demand-side behaviour commoditises the model layer in the same quarter that supply-side scarcity lets suppliers raise price. A poor place to own equity and an unreliable place to buy inputs, at the same time.
What the free tier now covers
Meta's Muse Glimmer is the specific fact that resets the COGS line: 30B parameters under Apache 2.0, running on a single 24GB RTX 4090 or a 32GB Apple Silicon Mac, sub-20GB at roughly 4-bit quantization for 0.2-1% accuracy loss, and about 233 tokens/second on an RTX 5090 using DFlash speculative decoding for roughly 3x throughput. It was trained around the agent loop itself, planning, tool calls, self-checking, failure recovery. Nous Research came at the same line from software, collapsing twelve Hermes browser tools into one and cutting token consumption 48-66% with no accuracy drop. Databricks has published its own token-cost optimisation methodology, which is the tell: cost-per-token engineering is table stakes, not differentiation.
| Lever | Cost effect | Verified? | Hard limit |
|---|---|---|---|
| Local 30B open weights | Variable per-token cost becomes amortized hardware | License and hardware specs verifiable; capability claim has no published benchmarks | Will not match frontier reasoning on hard turns |
| Tool-surface redesign | 48-66% fewer tokens per task | Reported with no accuracy loss | One-time refactor any competent team can copy |
| Hosted frontier API | Repricing upward, revenue share in play | Price direction reported, magnitude undisclosed | Capability ceiling is the reason you pay |
The honest caveat: a 30B model does not replace a frontier model on genuinely hard problems. Most agent turns are not hard problems. They are tool calls, retries and the self-checks between them, which is what this class of model was tuned for, and also where the token volume in every agent forecast was supposed to come from.
Where the two readings diverge
One reading says suppliers gained pricing power. The other says harness efficiency cuts revenue per task even as agent volumes climb. Both hold at once, price per token up, tokens per task down. A model vendor whose top line rests on agentic usage growth now needs a token-efficiency discount applied to the per-task assumption. An application company loses predictability in the variable-cost line in either direction, and it is the unpredictability rather than the level that breaks a five-year model.
The mirror image is the trade almost nobody is pricing. This is probably too clean, but zero-egress, zero-per-token local agents under a permissive license look like a structural advantage in healthcare, legal, defense and EU data-residency accounts, where cloud-API competitors cannot follow on price or on data handling. Two ways that goes wrong before the license does: the frontier gap widens faster than the harness improves, or enterprises decide dual-sourcing means two clouds rather than one cloud and one laptop. The unhedged tail risk is license durability: a business whose entire cost structure assumes permanent unrestricted access to one open-weight lineage is carrying regulatory risk it has not named.
If a holding's margin case improves only because tokens got cheaper, you own a market condition, not a business.
What to do
Commission a token-cost sensitivity across the ten largest AI-native holdings within 30 days, modeling gross margin at flat, +25% and +50% inference cost plus a supplier revenue-share scenario.
Require documented dual-sourcing and a model-abstraction layer as a diligence condition this quarter for any company where inference exceeds 15% of COGS.
Ask every AI portfolio CEO for a written local-inference and data-residency plan by month-end, then score which ones could sell into healthcare, legal, defense and EU-residency accounts.