The 80% Cut Is Rented Margin
Two independent supply-side datasets say the cost floor under your invoice rose while the rate card fell, which makes this a two-quarter window rather than a new baseline.
The spread is the instruction, not the discount
A platform PM opened the new rate card and backed out the pre-cut numbers. The decision was already sitting there. Luna and Terra used to sit roughly 2.5x apart ($1.00/$6.00 against $2.50/$15.00). They now sit exactly 10x apart on input and output both, per The Information's reporting. That is not a discount. It is a routing instruction. Anything left on the mid tier that a user cannot tell apart from the low tier is paid capability nobody sees.
An agent session consuming 50k input and 5k output tokens runs $0.16 on Terra against $0.016 on Luna. A million sessions a month is $160k versus $16k. Free-tier limits rationed on inference cost may not need the ration anymore.
Where the saving leaks back out
Then two numbers say unit price is the smallest term in the cost equation. Amazon burned $1.8M on a single routine coding task, 860% over budget, and found out only afterward through internal AI usage metrics. Stripe named the mechanism when describing Kai, the internal knowledge platform most of its employees use: most Kai sessions require many turns. Turns-to-resolution sets cost per resolved task. Price per token does not.
DeepSeek's rate card makes the point from the other side. V4-Flash 0731 charges $0.14 per million uncached input tokens and $0.0028 cached. No model choice repairs a 50x delta. An agent that serializes a fresh tool schema or a session ID into the top of its prompt pays the uncached rate on every request, forever. Artificial Analysis qualified its own Pareto-frontier claim on the near-99% cache-hit path, then publicly corrected a cache-hit calculation minutes after a screenshot of it circulated. That is a fair read on how few teams measure realized hit rate at all.
The floor under the invoice did not move
Here the sources genuinely disagree, and the disagreement is the intelligence. The Information reads falling API prices against inflating hardware as a reason to rent inference rather than self-host: Amazon lifted 2026 capex to $220B from $200B, which Andy Jassy attributed to memory chip costs, while Apple's CFO warned advanced chip constraints will increase significantly this quarter. a16z reads the supply side and finds the opposite for anyone modelling their own compute, with 12-month H100 contracts just under $2.50/GPU-hour, roughly 40% above November, Kalshi forwards near $2.78 and A100 spot merely stable rather than collapsing. a16z also dismantles the chart circulating as proof that AI demand is falling. Silicon Data's Token Cost Index measures token-spend intensity, so a mix shift toward cheaper tokens drags it down while B2B spend on Cursor, Anthropic and OpenAI rises in total and at the median across the four biggest-spending industries.
Read together the picture is coherent. The token price on the invoice is a competitive concession, granted three weeks after launch in response to customer bill shock, and it is reversible. The compute underneath is not getting cheaper on any product timeline.
A price cut you did not earn can be withdrawn. Cache-hit rate, turn count and spend ceilings are cost reductions you own.
So the repricing exercise has a second half most teams will skip. Re-score the shelved features at the new rate card, then re-run the survivors at flat and +25% compute. Two questions decide the sprint: does the feature still clear at +25%, and does its cost scale with turns or with tokens. Features that only clear at today's prices belong on a shelf list with a named trigger, not in next quarter's commitment.
What to do
Re-score every AI feature killed on unit economics since Q1 against the new low-tier rate card this sprint, then re-run the survivors at flat and +25% compute before scope locks.
Instrument realized cache-hit rate and turns-per-resolved-task on your two highest-volume AI surfaces this sprint, and report both next to activation and retention.
Ship hard per-job and per-workspace spend ceilings with a 50%-of-budget alert before any additional agent capability reaches GA.