Opus 5.5's Discount Is Proven Per Token, Unproven Per Task
What you can actually save depends on success rates that no public leaderboard can measure for your code. The eval you build this sprint decides which shelved features come back.
Only one cost term is settled
Your re-score needs one number: cost per successful task. That is tokens per attempt, times price, divided by success rate. Anthropic's release fixes the price term and shifts the token term. Simplifying AI worked the arithmetic. The per-token price fell 20%, so a net saving near 40% implies about 25% fewer tokens per task (0.8 × 0.75 ≈ 0.6). That is a derived figure, not an Anthropic disclosure. It still matters for your dashboards. A dashboard that compares models on dollars per million tokens will miss roughly half of this gain, and every future efficiency gain like it.
No vendor can give you the success-rate term. The SchrodingerRepo researchers found that on rewritten repos, 81.6–83.6% of the extra actions agents took were exploration, meaning the model searching through code it had never seen. That searching is what pushed input tokens up. Your private codebase triggers the same behavior, because no model has trained on it.
What unfamiliar code does to a business case
TheSequence ran an illustration with the paper's GPT-5.4-mini result. On a rewritten repo, input tokens rose 2.5x and first-try success fell from 46.8% to 35.6%. Together, that means roughly 3.3x the input-token cost per success. Even after a 40% price cut, you are near 2x what a leaderboard-based projection implies. The math mixes models and setups, so read it as a direction, not a forecast.
This risk doesn't hit all your planning documents equally. A cost baseline pulled from production already reflects your own code. A business case built from vendor demos or SWE-bench scores does not, so that is the one most likely to overstate ROI.
Both outlets say the headline figures are self-reported, but they differ on how fast to move. Simplifying AI treats Opus 5.5 as a model swap an existing Anthropic customer can make this week. It also suggests a routing pattern: default to Opus 5.5 and send only the hardest requests to Fable 5.1. That works because Anthropic claims Fable 5.1-level quality on “most tasks.” Simplifying AI adds that neither “most tasks” nor “typical workloads” is defined. TheSequence puts more weight on checking results against your own code before you commit.
A cost lever you control
Salesforce AI Research built JIT Mem, which stores raw records of past agent runs and assembles task-specific context only when a task needs it, instead of summarizing memory in advance. In the ALFWorld, WebShop and τ²-bench test environments, it cut input tokens 50.3–56.3% and steps 28.4–31.4%. It also beat memory systems that summarize in advance by up to 16.3 success points. Even an untrained Gemini curator scored 61.0, against 41.0 for SkillOS. These are benchmark environments, not production codebases. Still, JIT Mem targets the same exploration overhead the rewrite study exposed, and memory design is a decision your team owns.
The smart move
Build the eval before the business case. Draw 50–100 tasks from your own repos and workflows. Add variants that change how the code looks but not what it does, such as renamed namespaces and reordered files. They show how much a score depends on code the model has seen before. Then re-score the shelved backlog from those results, not from the launch post. Finally, clean up your own claims. Technical buyers can now easily challenge unqualified SWE-bench Verified scores in your PRDs or sales decks.
What to do
Build a 50–100 task private eval from your own repos and workflows this sprint. Include variants that preserve what the code does, and compare Opus 5.5 with your current model on success rate and tokens per successful task.
Re-score every AI feature you shelved for cost or latency using cost per successful task from that eval, and bring the top three into Q4 planning before the roadmap locks.
Audit PRDs, sales enablement and marketing pages this sprint for unqualified SWE-bench Verified claims. Add caveats or replace them with private-eval results.