Product & Strategy

The Product Desk

The Signal

Microsoft killed AI features across 81 products this week after customers called them

The dividing line: features that automate a task users already hate, producing output good enough to ship without editing, live. Everything else dies.

In Play

  1. Microsoft's 81-Product AI Pruning: The Kill Framework

    Microsoft killed Copilot in Gaming, Photos, Widgets, and Notepad while 365 Copilot grew 33% in paying users. The survivors automate weekly workflows (meeting transcription, email drafts). The dead added chat surfaces where nobody was asking questions. Nadella merged all Copilot under one EVP to enforce centralized quality gates.

    Ask Clarity
  2. Agent Payment Rails Ship: Stripe, Google Cloud, Coinbase Go Live

    Stripe shipped 280+ agentic commerce features including agent wallets. Google Cloud launched pay.sh with Solana for $0.001-$20 stablecoin micropayments. Coinbase shipped agentic.market. NFX's Project Deal proved agents close real transactions ($4K, 46% WTP). Per-seat pricing's expiration date moved from 'someday' to 'this cycle.'

    Ask Clarity
  3. AI Inference Margin Crisis: The 30-Point Gap Has a Fix

    AI features run at 50-60% gross margins vs. 80-90% for traditional SaaS. Reasoning models push per-query cost 10x higher. But vLLM delivers 2-4x speedups, prompt caching cuts 90% of repeated costs, and EAGLE decoding adds 45% speed. At 10M daily requests, a 30% optimization saves $20M+/year. Microsoft admits even free-model inference drags margins.

    Ask Clarity
  4. Verification > Generation: RAG Cliff + Legal Exposure

    RAG accuracy collapses from 90.7% at 5K docs to 50.6% at 500K — a coin flip at enterprise scale. Character.AI is being sued for fabricating a medical license number. Apple paid $250M for marketing undelivered AI features. Dario Amodei named verification as the Amdahl's Law bottleneck. Value is migrating from output generation to output trust.

    Ask Clarity
  5. GitHub at Zero Nines: AI Load Breaks Developer Infrastructure

    GitHub dropped to 86% uptime — 2-3 hours of daily degradation. AI agents drove 3.5x load growth in 2 years; GitHub revised capacity targets from 10x to 30x in 4 months. Mitchell Hashimoto (18-year user) left publicly. Competitors handle similar load without collapse. Any product shipping through GitHub Actions has release cadence coupled to daily outage windows.

    Ask Clarity

Deep Dives

Ship or Kill: Microsoft's $100B Pruning Experiment Gives You the Decision Grid

The Experiment, Concluded

A user opened Notepad this week to jot a phone number. The Copilot button was there. She did not press it. She has never pressed it. Microsoft shipped Copilot into 81 distinct product surfaces over 18 months and this week killed it in Gaming, Photos, Widgets, and Notepad, while 365 Copilot grew paying users 33% quarter-over-quarter. Pavan Davuluri's phrasing gives it away: users "want it to be better," not "want more of it." Coverage was never the product.

AI distribution is not AI value. The survivors live inside paid workflows and replace work the user was doing by hand. The dead added a chat surface to utility apps where no conversation was happening.

The Grid That Predicts Survival

Two axes fall out of the data. Axis 1: the AI replaces a task the user actively avoids (meeting transcription, email drafts), or it adds a layer to a task the user already does competently. Axis 2: the output has to be correct, or plausible is enough. The shippable cell is "replaces avoided task" where "plausible is enough." Every other cell is a demo with a roadmap ticket attached.

Why This Applies to 17 AI Features on a Roadmap

Meta's internal token-consumption leaderboard was gamed immediately. Engineers wrote scripts that burned millions of tokens doing nothing. Meta shut it down. The same pattern shows up in product dashboards. Teams count tokens consumed and sessions opened and call it adoption. The durable numbers are harder to collect: retention of AI-assisted workflows after 30 days, time-to-first-useful-output, percentage of AI output that ships to production without rewrite.

Multiple sources confirm the bottleneck has moved from engineering capacity to discovery and specification quality. When shipping takes an afternoon, shipping the wrong thing takes an afternoon too. Feature count goes up. Time-to-value does not.

The Margin Forcing Function

Microsoft gets OpenAI's technology at preferential rates and still admitted Copilot inference costs drag margins. Every AI feature carries an ongoing compute cost that scales with usage, not with value delivered. A feature with 5% engagement and 100% inference cost on every page load is burning money on 95% of impressions. Microsoft is not cutting features because users hate them. It is cutting features because each impression has a marginal cost and most impressions don't earn it back.


The Organizational Response

Nadella merged consumer and enterprise Copilot under a single EVP, Jacob Andreou. The diagnosis is governance, not product. When every product team independently bolted a chatbot onto its surface, nobody owned the holistic experience or the total inference bill. Teams with distributed AI feature ownership and no central quality gate are on the trajectory Microsoft was on six months ago.

What to do

  1. Map every AI feature in your product onto the 2x2 grid (replaces-avoided-task vs. adds-layer, and plausible-output vs. must-be-correct) by end of this sprint

  2. Pull 30-day retention and usage depth for each AI feature shipped in the last 90 days — replace token/session metrics on leadership dashboards

  3. Propose a centralized AI product owner or quality council to leadership using Microsoft's 81-product cautionary tale

  4. Kill or pause at least 3 AI features that show no usage lift after 4+ weeks in production

Your Next Buyer Has No Eyes: Agent Commerce Infrastructure Went Live This Week

Three Payment Rails, One Week

A product manager watching the agent-commerce space opened three tabs this week and found the same pattern shipping from different directions:

  • Stripe Sessions 2026: 280+ features organized around "agentic commerce" — agent wallets, autonomous transacting across fiat and stablecoin rails, AI checkout optimization
  • Google Cloud + Solana: pay.sh puts Gemini, BigQuery, and Vertex AI behind stablecoin micropayments at $0.001-$20 per call, no account creation required, MCP-server compatible with Claude
  • Coinbase: agentic.market with x402 protocol; Anchorage launched KYA-compliant Agentic Banking
When Google is willing to sell its own AI services on Solana stablecoin rails at a tenth of a cent per call, the interesting thing is not the crypto. It is the pricing model.

The Demand Signal Is Real

Anthropic's Project Deal put 69 participants through actual agent-to-agent negotiation. The agents moved $4,000 in actual transactions. 46% of participants said they would pay to keep using it. That is the number that matters — not demo traffic, not pilot NPS, willingness-to-pay on a workflow that already exists. NFX published a category thesis calling agentic marketplaces "bigger than SaaS" to 160K+ founders. Expect 50+ funded startups inside 12 months.

Per-Seat Pricing Has a Timer

HubSpot is pursuing full API parity with their UI so "agents can run on HubSpot, and agents can run HubSpot." Coming from a company with 200K+ customers, that sentence is the admission that the human dashboard is no longer the primary surface. ServiceNow shipped AI Control Tower with shadow agent discovery. Microsoft's Agent 365 is GA. The governance layer for agent-as-user is shipping faster than most products are designing for it.

The thing pitched as "AI pricing strategy" is usually a discount grid. The thing actually being done is re-pricing for a customer that calls an API ten thousand times in an hour and then goes quiet for a week. That customer is a terrible subscription and a natural micropayment customer. A product with an API and no agent-native billing path has a distribution channel that is closed on Monday.


The Standards War

StandardBackerBest ForDifferentiator
pay.shGoogle + SolanaAI developers75 providers, MCP/Claude native
x402 / agentic.marketCoinbaseConsumer appsExchange liquidity, brand trust
Agentic BankingAnchorageEnterprise/regulatedKYA identity, compliance

Developer adoption velocity decides this one, not spec elegance. pay.sh has early advantage with 75 launch providers and Claude compatibility.

What to do

  1. Map your product's transaction flow from an AI agent's perspective — identify every point requiring human presence (CAPTCHAs, email confirmations, visual browsing, seat-based auth)

  2. Model revenue impact if 30% of 'seats' become AI agents over 3 years — identify which pricing levers (API calls, outcomes delivered, data processed) supplement per-seat

  3. Evaluate pay.sh integration as a distribution channel for any API you ship — test whether per-request stablecoin payments unlock agent usage impossible under current billing

  4. Add 'agent-as-user' persona to your next user research cycle with specific questions about delegatable transactions

The 30-Point Margin Leak: An Inference Optimization Playbook You Can Execute in Weeks

The Problem Is Structural

A PM looked at her inference line last month and saw it had tripled while revenue doubled. She is not alone. AI features run at 50-60% gross margins against 80-90% for traditional SaaS, per BVP data. This is not a phase companies grow out of. COGS scales with usage in a way SaaS COGS never did. Every API call has a real marginal cost. Reasoning models make it worse. A single deep-reasoning response burns 100 turns' worth of tokens, pushing per-query cost up 10x even as per-token costs fall.

If the company with free models cannot make the economics work at scale, the smaller shops pretending they can are fooling themselves.

The Optimization Stack Is Production-Ready

A year ago these techniques were conference talks and research papers. Today they ship as infrastructure:

  1. Prompt caching: Anthropic's technique cuts 90% of costs on repeated system prompts. vLLM plus Mooncake hit 92.2% cache hit rates on agentic workloads, up from 1.7%, with 3.8x throughput and 46x lower time-to-first-token.
  2. Model routing: Route 80% of queries to fast, cheap models. Reserve expensive reasoning for the cases users will pay a premium for. Google's GKE Inference Gateway routes to warm caches automatically.
  3. vLLM + speculative decoding: vLLM delivers 2-4x speedups from a backend switch. EAGLE speculative decoding adds 45% on top. Both are in SGLang and TensorRT-LLM already.
  4. Full-stack stacking: Yandex published a blueprint combining quantization, EAGLE3, KV cache reuse, and parallelization for 5.8x total speedup, cutting token generation from 140ms to 67ms.

The Dollar Math

At $0.02 per inference and 10M daily requests, the inference line runs $70M a year. A 30% reduction, which any one of the techniques above can deliver on its own, saves $20M+ annually. That is not an engineering efficiency metric. It is a P&L line larger than most features sitting on the roadmap this quarter.

The Anthropic Advisor Pattern

Anthropic shipped an advisor strategy this week. Sonnet runs as primary inference and calls Opus on-demand when the problem warrants it. The claim is frontier-model quality at 5x lower cost. The old binary of expensive-and-good versus cheap-and-worse no longer describes the Pareto frontier. OpenAI claims 1,000x cost reduction over 14 months through stacked optimizations.


The PM Decision

Inference optimization is not a platform ticket to defer to next quarter. It is a product commitment, and the PM owns the margin number the same way they own retention and time-to-value. The forcing function is straightforward. Put the annualized dollar value of a 30% inference cost reduction on one side of the page. Put the expected revenue contribution of Feature X on the other. At scale, optimization almost always wins the sprint.

What to do

  1. Build a per-feature inference cost model mapping each AI capability to actual cost-per-use — present to finance as 'AI ROI dashboard' within 2 weeks

  2. Commission a 2-week infra spike on prompt caching and vLLM adoption for your highest-volume AI endpoints

  3. Add 'inference cost impact' as a required field in your PRD template for any AI-powered capability

  4. Implement tiered model routing: fast/cheap models for 80% of queries, expensive reasoning only where users demonstrably value it

Verification Is Where Value Migrates: The RAG Cliff, Credential Lawsuits, and the $250M Apple Precedent

The RAG Scaling Cliff, Quantified

A PM ships a knowledge assistant into a 5K-document pilot and it works. Retrieval feels crisp. Legal signs off. Six months later the corpus hits 500K documents and the complaints start. That curve has a number on it now. Onyx's EnterpriseRAG-Bench measures vector search accuracy falling from 90.7% at 5K documents to 50.6% at 500K. At enterprise scale, the retrieval layer is a coin flip. The mechanism is not mysterious. At 5K docs, 3-5 documents touch any given topic and top-k retrieval lands the right ones. At 500K, 40-60 documents sit in the same embedding neighborhood and the system cannot tell them apart.

BM25 keyword search degrades more gracefully (85.8% → 68.4%), which reframes hybrid retrieval P0 infrastructure, not a nice-to-have. Adding BM25 reranking is typically 1-2 sprints for roughly 18 percentage points of accuracy recovery. Knowledge-graph architectures like Rowboat (13K+ GitHub stars) hold query cost flat as the corpus grows. That is a different scaling curve, not a faster version of the same one.

Legal Exposure From Unverified Output

Three legal developments landed in the same week:

  • Character.AI: Pennsylvania sued after the chatbot fabricated a specific medical license number and claimed to be a licensed psychiatrist. A novel consumer-protection theory any state AG can copy.
  • Apple: Settled for $250M over AI features marketed before delivery. 37M devices, $25-$95 per claim. First enforceable precedent coupling GTM timeline to engineering delivery.
  • Connecticut SB5: Automated decision-making is codified as not a defense to discrimination. Deployers own discriminatory outcomes whether or not AI made the call.
As generation gets faster, verification becomes the serial bottleneck. Value is moving from 'AI writes it' to 'can I trust what AI wrote.'

The Product Implication

Dario Amodei invoked Amdahl's Law on AI coding: generation speed is irrelevant if a human still spends 40 minutes reviewing a 400-line diff trying to decide what to trust. Cognition's Devin Review targets that gap directly. The feature the team is pitched on is generation throughput. The feature users actually need two quarters out is the one that cuts verification time: confidence scoring, automated review, audit trails, formal verification. Generation is commoditizing. Verification is not.


The Design Response

Here is the 2x2 worth drawing on a whiteboard this sprint. One axis: corpus size above or below 100K documents. Other axis: does the product act on the output, or surface it for human approval. Products in the large-corpus-plus-act-without-checking cell face a 50.6% accuracy rate that is functionally a lawsuit. The fix hierarchy is unambiguous. Ship hybrid retrieval with BM25 reranking this sprint. Put reranker investment ahead of the next embedding model upgrade. Get a knowledge-graph-based architecture on the roadmap before the corpus exceeds 100K docs. For conversational AI, the minimum viable guardrail is an output classifier that detects credential claims, professional titles, and license-number patterns before the response leaves the server. That is the Character.AI case, caught at the edge.

What to do

  1. Run your RAG pipeline against EnterpriseRAG-Bench at projected 12-month corpus size — not pilot size — before next enterprise deal closes

  2. Add hybrid retrieval (BM25 + vector search) to your retrieval architecture if not already present — 1-2 sprint investment for 18-point accuracy recovery

  3. Audit all customer-facing AI claims — marketing pages, app store descriptions, sales decks — for features marketed but not fully delivered, using Apple's $250M settlement as the legal bar

  4. Add an output classifier detecting credential claims, professional titles, and license-number patterns to any conversational AI feature before next release

The bottom line

Microsoft spent $100B shipping AI into 81 products and just proved that AI distribution is not AI value — the features that survived automate hated weekly tasks with good-enough output, everything else burned inference cost for zero retention. Your roadmap has the same split: run the 2x2 audit this sprint, kill the features in the wrong cell before they become your own 'functionally useless' story, and redirect that inference budget into the optimization stack (vLLM + prompt caching recovers 20-30 margin points in weeks, not quarters) and the verification layer where defensible value is migrating as generation commoditizes.