The Cheapest Margin Win This Sprint Is Your Own Tool List
Deleting tool definitions beat renegotiating a contract, and it arrives while two model providers signal they want more money, not less.
Why a tool list is a cost line
An engineer registers the twelfth tool on an agent, runs it, watches it work, and ships. Nothing broke. Nothing visibly breaks when the list grows, which is why nobody owns deletion, and why agent tool surface accumulates the way feature flags accumulate in an aging codebase. The bill arrives on every turn instead. Twelve schemas, descriptions and parameter lists get serialized into the context of every single call, used or not, and multi-turn is where agents actually live. That is what makes Nous Research's swap to a single higher-level tool driven by browser-use CLI 3.0 worth reading twice: a margin win with no vendor dependency and a clean success metric, since the model did not change and the contract did not change. Separate the pitch from the work, though. Consolidating twelve tools into one moves branching logic out of the model and into a CLI someone on the team now maintains. A fixed eval suite is the guardrail. A demo that looks fine is not.
Where the evidence agrees, and where it thins
The direction has corroboration. Stagehand v4's hardening priorities are context management, self-healing actions and iframe support, which is a different team arriving at the same conclusion: context discipline is the reliability lever and the cost lever at the same time. Agent "skills" — named, persisted multi-step process definitions with explicit triggers and stop conditions — are displacing ad hoc prompting on the same logic. A bounded spec costs less per run and fails more predictably than an open prompt.
The magnitude is weaker than the direction. The 48-66% figure is one team, one harness, one eval set, self-reported. That is a hypothesis with a measurement protocol attached. It is not a number for a board slide.
| Cost lever | Who owns it | Time to a result | Dependency |
|---|---|---|---|
| Tool-surface consolidation | Product plus engineering | Days | None |
| Model right-sizing | Engineering | Weeks | Eval coverage |
| Prompt and context caching | Engineering | Weeks | Provider feature support |
| Price renegotiation | Procurement | Quarters | Vendor's willingness |
Why the clock is running
Next quarter's planning is too late, because input price is moving the wrong way. DeepSeek has notified users of a significant price increase. Reuters reports Alibaba is considering asking for revenue share, a structure that lands on the P&L rather than sitting neatly in a COGS line, and one most product pricing cannot absorb. Benedict Evans frames the current market as a supply crunch in which labs can name their price, with no clear answer on when supply, demand, capacity and capex re-equilibrate.
So the "inference gets cheaper forever" assumption is retired for the next several quarters. The forcing function is arithmetic. Run each AI feature at 1x, 1.5x and 2x current token cost. Anything gross-margin negative at 1.5x needs a named mitigation before the next pricing commitment: caching, model right-sizing, a usage cap, or a price change. Named now, it is a roadmap decision. Named after a price is published, it is a customer conversation.
The one margin lever that needs no vendor's permission is the list of tools you handed the agent.
What to do
Count the tools exposed per agent across every production agentic feature by the end of next week, merge overlapping tools, and measure tokens per completed task before and after against a frozen eval set
Re-model each AI feature's gross margin at 1x, 1.5x and 2x current token cost this sprint and attach a named mitigation to anything that turns margin-negative at 1.5x
Add tokens-per-completed-task with per-feature attribution to the weekly product dashboard within the quarter