Opus 5.5's Discount Goes to Whoever Customers Can't Swap Out
A cut this deep only improves margins at companies with pricing power, and open-source routing tools have made that power harder for most AI apps to claim.
The 40% is two levers, and only one is guaranteed
Simplifying AI splits Anthropic's number into two parts: a 20% cut in per-token price and roughly 25% fewer tokens per task. Together they bring cost to about 0.6 of the old level. The price cut applies to every call. The token saving depends on the work being done, and Anthropic measured it on “typical” workloads. A portfolio company running long agent loops on unusual code may capture less. Three claims come from Anthropic's own measurements, not independent evaluations: the cost figures, the more-than-30% speed gain, and the claim that Opus 5.5 matches Claude Fable 5.1 on most tasks.
If you have exposure to the model layer itself, the arithmetic runs the other way. Simplifying AI calculates that each workload moving from Opus 5 needs about 1.67x more volume to keep Anthropic's revenue flat. Anthropic is pricing a Fable-class model below its own flagship. That is deliberate self-cannibalization aimed at winning share in coding and agents.
Why most of the saving won't stay with the companies you back
Every rival of a Claude-based app got the same cheaper model on the same day. It went live on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry. TheSequence's read is that the saving stays with an application only where the company has pricing power. Otherwise competition hands it to customers. Simplifying AI expects that pass-through within a few quarters.
The same week also lowered the cost of switching. HarnessRouter puts Claude Code behind one interface alongside Codex and DeepSeek's harness, so changing an agent's backend becomes a configuration edit. CopilotKit released OpenMuse, a free, MIT-licensed clone of Meta's new Muse agent. Architecture Notes reports that Google's ax is trying to set the convention for how agent sandboxes are defined. With switching this cheap, a company's claim to keep the saving has to rest on its workflow data, contracts and distribution, not its model access.
| Release | Price move | Basis for the claim | Likely beneficiary |
|---|---|---|---|
| Claude Opus 5.5 (Anthropic) | ~40% lower net cost than Opus 5; $4/$20 per million tokens | Anthropic's own measurements | Apps with pricing power |
| MiMo-V2.6-Pro (Xiaomi) | Free weights, MIT license | Third-party index: 46.32, top open-weight score | Managed hosts able to serve >1T parameters |
| Hy Image 3.5 (Tencent) | ~$0.024 per image vs ~$0.12 implied for Google's Nano Banana Pro | Tencent's own designers | Image-heavy apps, if quality holds |
The open-weight price floor needs one correction: free weights are not free inference. MiMo uses about 42B parameters per token, which keeps compute per token manageable. But holding more than 1T total parameters in memory takes multi-GPU serving that most enterprises won't run themselves, so that demand flows to managed inference hosts. Simplifying AI flags a possible limit, which it labels as its own inference: Chinese-origin weights may face procurement friction in regulated sectors. The source also does not report how MiMo's score compares with closed frontier models.
For image-heavy apps, Simplifying AI's illustration is concrete. At 10M images a month, an app would pay about $240K on Hy Image 3.5 versus about $1.2M at Nano Banana Pro's implied price. That saving only materializes if Tencent's self-reported quality holds up under independent testing.
The smart move
Treat the cut as a pricing decision each portfolio company must make on purpose, not a margin gain to book automatically. A company that plans to keep the saving should be able to say why its customers cannot route around it. A company that plans to pass it through should show the volume or share it expects in return. Either answer is defensible. Without one, a competitor makes the decision for it.
What to do
Ask every Claude-dependent portfolio company this week to re-run inference COGS on Opus 5.5 using its own traffic, and to bring an explicit keep-or-pass-through pricing decision to its next board meeting.
Commission a revenue-sensitivity update this quarter for any model-layer exposure, assuming workloads migrating from Opus 5 need about 1.67x the volume to keep revenue flat.
Map managed-inference hosts that can serve 1T-parameter open-weight models in US and EU regions before your next AI-infrastructure IC, including each host's exposure to procurement friction over Chinese-origin weights.