Two Companies Now Set the Token Price, and Neither Needs It to Be Profitable
Every AI application mark in the book was underwritten on an input price that turns out to be a strategic choice made by parties earning their return somewhere else.
Opposite mechanics, identical consequence
Meta's motive is compute absorption: Muse Code reportedly earns revenue from AI while the company spends heavily on data centers, so the contributor tier isn't priced against a gross margin at all, and Alexandr Wang naming Claude and Codex as the targets tells you the profit motive sits in a different P&L. DeepSeek runs it backwards, or rather runs the more interesting version backwards: demand for V4-Flash overwhelming finite capacity, at a firm whose founder told investors profit was not the priority. One discounts because tokens are a byproduct, the other raises because tokens are rationed. Neither is a cost curve, and that is the whole finding.
The buyer side already behaves this way
Microsoft now gives each internal department a limited pool of AI tokens to cap GitHub Copilot inference cost. Amazon, Google and Microsoft are committing roughly $600B of capex, and AWS says publicly it won't have enough capacity to meet demand through 2027. The largest AI spender on earth put a budget ceiling on its own flagship AI product in the quarter it bought more capacity than anyone. Infrastructure scarcity is contracted and underwritable; application-layer consumption has entered budget-committee discipline.
The application layer has said it out loud twice. Canva's COO: serving free users used to be "very low" cost, and with AI "those costs became much higher. The unit economics changed." Figma's CFO, on products still in beta: "we bear the cost of inference without offsetting consumption revenue". Figma reported a $370M quarter, guided growth down from 48% to 36%, and fell roughly 15% in a session. An unpriced beta feature is an open-ended liability with no end date.
What it costs when nobody meters the loop
a16z's Yoko Li ran Anthropic's own widely-copied loop example against a page whose achievable Lighthouse score she capped near 89. The first $1.40 moved the score from 26 to 89. The remaining $2.84, 67% of the $4.24 bill, bought exactly zero points, with per-turn cost rising as the transcript grew. A Haiku evaluator quietly billed $0.67 of its own, and when Claude correctly declared the goal impossible around try five, that evaluator bounced the verdict back 14 times. Nobody saw any of it until the trace was read afterward. The returns curve underneath is logarithmic: one web-agent benchmark went from 38.8% to 43.2% success between one sample and ten, then bought 0.2 additional points for double the tokens at twenty. n=1, benchmarks unnamed, directional rather than audit-grade.
Where the sources diverge
One reading treats funded pain as a sourcing lane: distillation, routers, semantic caching, cost observability. Possibly right. But both companies with the most acute pain chose to build in-house, and Canva's proprietary model wasn't ready in time, which is precisely what forced the rollout pause and the guidance cut. Either the vendors get bought into that gap, or the incumbents ship their own models a quarter late and eat the guidance. The second, held loosely: model timelines now sit on the critical path to revenue guidance, and hours spent making inference cheap are hours not spent shipping the feature meant to pay for it. Treat the vendor claim that in-house models are cheaper and faster than frontier alternatives as unaudited.
An input price set by a company that does not need margin on it is not a cost curve; it is a policy, and policies revert on someone else's schedule.
What to do
Commission a +30% inference-cost stress test across every AI application holding this month, flagging any company whose gross margin breaks below 55% or whose runway breaches inside 12 months.
Re-underwrite PLG and freemium marks on gross profit rather than ARR multiples before Q4 marks close, requiring cost per free monthly active user, cost per AI action, and margin at 3x adoption.
Add two gates to every new AI application term sheet this quarter: a dated metering commitment for unpriced beta features, and a demonstrated 30-day path to a second inference provider or open weights.