The Compute Pricing-Power Inversion: Five Simultaneous Attacks on the Nvidia Trade
What Changed
This is not one story. The more interesting version is that five independent signals converging in the same week all lean on the same assumption sitting under every GPU-dependent position: that Nvidia's pricing power holds. Alone they are data points. Together they are harder to mark around.
- Nvidia's revenue-share land grab: Nvidia now takes product revenue share and equity from AI startups in exchange for compute access. Startups accept to get around their own liquidity problems. Nvidia raised $20B in debt in the same breath, which means it is funding its own demand while collecting rent from customers who cannot afford to say no.
- Anthropic/Samsung custom silicon talks: The clearest confirmation yet that the largest buyers will vertically integrate. It mirrors the hyperscaler ASIC moves, except it comes from a model lab rather than a cloud provider.
- AMD MI355X posts >2x cost-efficiency: 2,626 tok/s/node on GLM-5.2, more than double comparable Nvidia setups. The software gap, the thing that kept the whole thesis intact for years, is narrowing at the workload level.
- 95%+ Grace-Blackwell undeployed: More than 18 months after first shipment, the vast majority sits unactivated. Price it in: when that capacity switches on, it meets a market that has been paying a scarcity premium, and premiums do not survive supply arriving in bulk.
- Nvidia Kyber rack delayed to 2028: A manufacturing failure on the circuit boards, and cloud customers rejected the two-rack workaround as too costly. That is a 12+ month gap for AMD and Google to walk through.
Why This Is Different From Monday's Coverage
Monday's briefing flagged Meta's managed-inference entry hitting the neoclouds. Today is structurally different, because the incumbent itself is making defensive moves. A monopolist that starts taxing your revenue instead of just selling you hardware is telling you the customers have gotten big enough, and annoyed enough, to build around it. The Anthropic/Samsung talks are that motive made concrete.
When Nvidia takes equity for compute and its biggest customer designs its own chips, the lever it uses to reprice hardware upward each cycle gets shorter.
The Contradiction Worth Noting
Foxconn's $79B quarter, up 40% YoY, confirms AI-server demand is still expanding. The tension is that demand grows while pricing power compresses. That is the ordinary late-cycle shape, revenue growth papering over margin erosion. Marks built on volume can miss the structural margin compression loading underneath them.
Inference-cost compression is accelerating from multiple vectors:
- OpenAI halved inference costs this quarter
- Alibaba's framework cut tokens 884K→1,160, a 99.87% reduction
- Condense cuts coding-agent bills 72%; pxpipe cuts 70%
- A 35B-parameter model now matches 1T models via horizon-scaling
Caveat: the 95% undeployed figure conflates 'delivered but idle' with 'not yet delivered,' and the exact supply timing is genuinely uncertain. The timing call here is probably wrong. The direction is harder to argue with, and the scarcity premium is on borrowed time.
What to do
Audit every compute-heavy portfolio company for existing or pending Nvidia rev-share/equity-for-compute agreements and model the dilution into your next mark
Build a custom AI silicon deal map by end of July — design services, packaging, ASIC challengers positioned to capture Anthropic-style vertical integration demand
Stress-test portfolio gross margins under a 30-50% inference-cost decline scenario and reclassify holdings by margin resilience
Reprice neocloud/GPU-cloud positions against a 5-15% Nvidia revenue-share take on cloud revenue