Product & Strategy

The Product Desk

The Signal

Blue-chip CFOs told UBS they want software spend down 30% because AI can rebuild it.

A procurement lead reopens the renewal thread and asks for active seats, not licensed ones. Seat-based licenses are being de-rated while consumption models get paid up, and IBM's mainframe slide shows that substitution already landing in reported results. Some of this is negotiating posture that fades once budgets reset. The packaging decision does not wait for the answer. What sits on the desk this quarter is whether a metered tier ships before the renewal date, not whether the feature set is good enough to defend the seat count.

In Play

  1. Buy-Versus-Build Repricing Hits Your Renewal Book

    UBS software lead Karl Keirstead says blue-chip executives are telling him directly that they want spend with software vendors down 30% over the next three years, because AI models are now good enough to custom-build alternatives, per The Information's briefing. For you that pressure arrives at renewal as a packaging question, not a feature-quality question. IBM's shares fell on declining mainframe purchases as customers redirected budget to AI — the same substitution showing up in reported results.

    Ask Clarity
    Try
  2. Agentic Engineering Finally Has A Public Denominator

    Zalando published rollout numbers for agentic engineering: 33% of pull requests now auto-approved as low risk, PR lead times down 20–40%, across more than 250 engineering teams, per The ML Engineer's issue 400. That gives you the first non-vendor benchmark to cite in an AI-tooling PRD — and the counter-metrics too. Zalando reports cyclomatic complexity, a measure of how tangled code paths are, spiking as adoption grows, with commit messages ballooning to 5,000 characters.

    Ask Clarity
    Try
  3. Deployment Actually Runs On Small Models

    Hugging Face's state of open models report finds models under 1B parameters account for 83% of all-time downloads, while models above 100B account for just 1%, as relayed in The ML Engineer's issue 400. Your default model tier for volume and latency-sensitive features is probably larger than your users need. The report only sees Hugging Face's own hub, so triangulate with your inference telemetry before locking a tiering decision — and note Chinese labs set the monthly open-model size ceiling, from 754B up to 2.78 trillion parameters.

    Ask Clarity
    Try
  4. A Court Can Revoke Your Category Noun

    The UK Supreme Court barred Oatly from using the word "milk" in February 2026, and in March a Swiss court struck down Alpro's "Shhh…This Is Not M*lk" as potentially misleading, per Morning Brew. Underneath the label fight, dairy-free retail volumes fell more than 5% for three consecutive years on Circana data, while dairy milk held strong on the protein trend. If your one-liner needs a competitor's noun to make sense, you are renting the category — and the incumbent can take the benefit back first.

    Ask Clarity
    Try
  5. The Capability Pricing Floor, Not The 2031 Date

    Peter Diamandis published a named-date forecast on August 16, 2026: by 2031, most people get access to what a $1M-a-year household buys today across eight categories, from education to transport. The usable part is the price anchors — private tuition at $40–60K a year against free AI tutoring measured on Bloom's 2-sigma effect size, and robotaxi service asserted at $50 a month. The date needs four separate technology curves to mature together, so treat it as pressure on willingness-to-pay rather than a planning input.

    Ask Clarity
    Try

Deep Dives

The 30% Cut Lands On Your Packaging Before Your Product

A CFO under a reduction mandate cancels seat-based line items first; usage-linked ones shrink instead — and which one you sell is settled long before the renewal call.

What the tape is de-rating, precisely

A procurement lead sorted her renewal calendar this week by contract type instead of by vendor, and the list rearranged itself into winners and casualties. That sort is the read. The market is not de-rating software. It is de-rating seat-based licenses and legacy infrastructure while paying up for consumption and AI-native models. The Information's briefing lines the evidence up cleanly enough to paste into a pricing review.

CompanyModel archetypeMarket verdictWhat you take from it
AtlassianCloud migration, expanding usageStock soared on cloud growthA credible consumption plus migration story still earns a premium
PalantirAI platform and deploymentContinued scorching revenue gainsIt is absorbing the AI-platform budget incumbents are losing
AnthropicModel API consumptionRevenue up roughly 14x year over year in Q2Company-provided figure with no absolute base disclosed: slope, not scale
IBMLegacy hardware and licensesShares crashed on falling mainframe purchasesBudget substitution showing up in audited results, not surveys

Workday is the case worth studying because the market cannot price it. Shares spiked 19% Thursday on the Reuters report that Silver Lake is in talks to acquire it, gave back nearly 4% Friday, and now sit 13% above the pre-report level. Up 55% since late June, still down 7% year to date despite $2.8B in free cash flow. Reuters carried no purchase price, so the buyout multiple everyone wants as an anchor stays unset. What teams tell themselves is that cash generation protects the price. What the tape says is that it protects the company and not the packaging, and nobody has repackaged yet.


The dated catalyst that costs nothing

Salesforce reports in roughly two weeks, with analysts polled by Refinitiv expecting revenue up more than 10% to $11 billion. Treat that call as primary research on the question being litigated in your own building: does AI arrive as incremental revenue for an incumbent, or does it only defend the base? Pull three disclosures off the transcript. Attach rate. Pricing structure. Whether AI revenue reads as incremental or substitutive. Those three go into the AI SKU business case before someone else challenges the projections with worse data.


Discount the number honestly

Two caveats belong on the same slide as the 30%. It is stated intent, not realized spend, and enterprises say this in every downturn. And KeyBanc's Jackson Ader attributes part of the recent software bounce to technical rotation, money leaving high-momentum names for the relative losers of late, which means the recovery may have no fundamental floor under it. What is different this cycle is that mainframe decline puts substitution into reported financials rather than survey responses. Plan for the intent being half true. Half of that mandate is still a repricing event across the whole renewal book.


One mechanism, three markets

This is not only a software story, which is why it should change sequencing rather than talking points. Diamandis's abundance forecast and the plant-based milk collapse describe the same mechanism from different angles. The capability a company charges for gets absorbed by a free model, by an incumbent's feature list, or by the buyer's own engineers, and the price follows within a cycle or two. The diagnostic fits on two axes: is the thing you charge for a capability or an outcome, and can the buyer obtain it without you next year. What survives in the first column is the defensible layer. Proprietary data. Integration depth. An audit and compliance surface. A measured outcome someone will sign for. Feature parity is not a defense against a buyer with a cut mandate.

A seat-based line item is the easiest thing a CFO under a cut mandate can terminate; a usage-linked contract shrinks instead of dying.

What to do

  1. Score every revenue-weighted feature 1-5 this sprint on how credibly a customer's own team could rebuild it with a current-generation model plus their own data, and ship the ranked list to your exec team with a named moat for everything at 4 or above.

  2. Draft a consumption or outcome-linked pricing variant for your top revenue line this quarter and pressure-test it with five existing customers ahead of their renewal dates.

  3. Mine Salesforce's late-August results within 48 hours of the call for AI attach rate, pricing structure, and whether AI revenue is incremental or substitutive.

Zalando Published The Agentic ROI Numbers Your PRD Was Missing

The copyable asset is not the gateway or the model mix — it is instrumentation good enough to make adoption legible, and candid enough to disclose the code-health regression underneath.

The layer that is actually copyable

An engineer at Zalando opens a chat window, picks a model, and never has to learn which vendor answered. That is the whole trick, and it is the unglamorous part. Zalando routes OpenAI, AWS Bedrock and Google Vertex models through an in-house gateway called zLLM, built on the open-source LiteLLM project, serving roughly 2,000 monthly active users and growing. One gateway produces one denominator. Active users, teams onboarded, and spend all land in the same table. Without that layer, an ROI slide is a narrative. There is a second-order effect worth naming alongside it: gateway monthly actives is simultaneously the adoption proxy and the inference-spend curve, so it grows with success rather than with budget.

The enablement machinery is equally copyable and almost never funded: weekly guild sessions, monthly trainings at 120-150 participants, hackathons, and scoped agentic pilots as the on-ramp. Scope is a political variable, not only an operational one. Once an org-wide number covering more than 250 engineering teams is public, a 5% internal pilot stops reading as progress to an exec team.


The counter-metrics are what make the story credible

Zalando published the parts that do not flatter it. That is the part to copy hardest. Cyclomatic complexity, a standard measure of how many tangled branching paths a piece of code has, is spiking as agent adoption grows. Commit messages are ballooning to 5,000 characters, which is a comprehension tax on reviewers and a leading indicator of review fatigue. There is explicit friction around tool migration and transformation too. The velocity gain is real. It is partly financed by code comprehensibility.

The forcing function is a single rule: every throughput metric on an AI-tooling dashboard needs a health counter-metric beside it, instrumented before rollout. Complexity trend, commit length, review rework rate. Debt discovered after adoption is discovered structurally.


Where the auto-approval number concentrates risk

Auto-approving a third of pull requests as low risk hands a governance decision that used to belong to a human reviewer over to a model-based risk classifier. That is a reasonable trade, and the two conditions sit in the same paragraph as the trade: name the risk classes that never route to it (authentication, payments, data-deletion paths, schema migrations), and have someone review the classifier's false-negative rate on a fixed cadence. Nobody publishes that number yet. The team holding it internally before an incident asks for it is the team that gets to keep the trade.


Throughput is the easy metric; effect size is the hard one

Set these numbers against the measurement standard emerging elsewhere. Diamandis's forecast benchmarks AI tutoring on Bloom's 2-sigma effect size, which is an outcome delta rather than an activity count. Zalando measures activity and speed. That is the correct starting point and not the finish line. The pattern running through today's material holds: a claim with a denominator survives scrutiny, and a ratio without one gets discounted the moment a finance partner reads it. Next quarter's agentic-tooling budget review will be won or lost on which of those two is in hand.

Caveat to carry with the numbers: this is a single organization's account, relayed by a practitioner who says he was personally part of the transformation. Treat it as a credible reference point, not an industry baseline.

Zalando's copyable asset is not its gateway or its models. It is that it can produce a numerator and a denominator on demand.

What to do

  1. Copy the four throughput metrics — auto-approval rate, PR lead time, gateway monthly actives, teams onboarded — into your AI-tooling dashboard this sprint, and pair each one with a code-health counter-metric.

  2. Define the risk classes that never route to model-based auto-approval before you raise your auto-approval target, and put the classifier's false-negative rate on a monthly review.

  3. Add the gateway inference-spend curve to your next planning cycle as a function of active-user growth, not a flat line item.

Oatly Lost The Word "Milk" After Dairy Took Its Only Real Benefit

Two courts revoked a challenger category's vocabulary within six weeks, but volumes had already been falling for three years — the naming loss was the symptom, not the cause.

Three markets, three legal product names

A shopper reaches for the oat carton and scans the front panel for the word she already uses at home. Regulators in three markets disagree about whether that word is allowed to be there, which makes product nomenclature a jurisdictional compliance surface rather than a brand decision. Same carton, three legal name constraints, and naming now sits in the launch-readiness gate next to data residency.

MarketStance on "milk" for plant-basedTriggerProduct consequence
United StatesPermitted; the FDA holds consumers already understand nut milks are not dairyStanding FDA positionNo renaming; clear labeling suffices
United KingdomProhibitedSupreme Court ruling, February 2026Packaging and marketing rewrite, new descriptor vocabulary
SwitzerlandProhibited, including ironic negationCourt ruling, March 2026Full product-name change and brand equity reset

The sequencing is the lesson. Alpro named the product "Shhh…This Is Not M*lk", a self-aware negation, and a court still ruled it potentially misleading. The category was already losing before either ruling. Dairy-free retail volumes fell more than 5% for three consecutive years on Circana data while dairy milk held, because consumers went hunting for protein. The incumbent did not win the label case first. It took the health halo, and the label case arrived afterward to remove the challenger's vocabulary.


What this turns into on a roadmap

Two exercises, both doable this week. First, delete every competitor and incumbent-category noun from the one-liner and check whether it still communicates value. If it collapses, what exists is not a position but a comparison, and comparisons are revocable by a court, a rebrand, or a search-ranking change. Second, name the single differentiator the largest incumbent could ship as a feature in one release note, then fund something they structurally cannot replicate: proprietary data, workflow embedding, distribution, or an outcome someone is accountable for. Diamandis arrives at the same idea from the other end when he concedes the skills layer goes free while judgment and accountability stay scarce. That is a moat statement dressed up as parenting advice.


Premium decoupled from trade-down, but do not over-read it

The incumbent is riding a premiumization tailwind that contradicts the trade-down narrative. US butter consumption hit a record 6.8 lbs per person in 2024 on USDA data, next to $60-per-pound Vermont farm butter, $15 ice cream pints, and third-wave frozen yogurt at up to $30. One Paris retailer sold 19 tons of butter, up 300% from 2023, largely to American tourists. That supports a test of a genuinely expensive top tier alongside a low-friction entry point, with hard scrutiny of whether the middle tier earns its existence. It does not support repricing the whole ladder: pandemic baking habits, flight from ultra-processed fats, and grocery inflation make this a lipstick-index pattern, and lipstick indices unwind when discretionary spend actually contracts.


The asset nobody could rebuild

"Got Milk?" reached an estimated 80% of US consumers on any given day and was retired in 2014. After that the category leaned on episodic virality: the 2020 #GotMilkChallenge with Katie Ledecky and Tony Hawk, then a celebrity chug at a 2025 award show. High-recall assets are cheap to reactivate and effectively impossible to rebuild from scratch. That inventory deserves an audit before a legacy feature or product name gets sunset in a rationalization pass.

If deleting your competitor's name breaks your positioning, you do not own a category — you are renting one, and the landlord can evict you.

What to do

  1. Run the noun-deletion test on your product one-liner this sprint: strip every competitor and incumbent-category term, and rewrite anything that stops communicating value.

  2. Commit one roadmap item this quarter to a benefit with a real replication barrier — proprietary data, workflow embedding, or distribution — and name the feature you are defunding to pay for it.

  3. Add jurisdictional naming and claims review to your international launch-readiness gate before your next UK or EU release.

The bottom line

Today's stories rhyme on a point none of them states outright: buyers have stopped paying for capability and started paying for proof that capability produced an outcome, plus a contract shape that survives a cut mandate. That retires the assumption that a differentiated feature list defends a price. Pricing power now sits with whoever holds a measured outcome and a benefit the customer's own engineers structurally cannot reproduce. Write one sentence this week naming the outcome your largest revenue line provably delivers — if it needs a competitor's noun or a feature list, that is your next roadmap problem.