Ship or Kill: Microsoft's $100B Pruning Experiment Gives You the Decision Grid
The Experiment, Concluded
A user opened Notepad this week to jot a phone number. The Copilot button was there. She did not press it. She has never pressed it. Microsoft shipped Copilot into 81 distinct product surfaces over 18 months and this week killed it in Gaming, Photos, Widgets, and Notepad, while 365 Copilot grew paying users 33% quarter-over-quarter. Pavan Davuluri's phrasing gives it away: users "want it to be better," not "want more of it." Coverage was never the product.
AI distribution is not AI value. The survivors live inside paid workflows and replace work the user was doing by hand. The dead added a chat surface to utility apps where no conversation was happening.
The Grid That Predicts Survival
Two axes fall out of the data. Axis 1: the AI replaces a task the user actively avoids (meeting transcription, email drafts), or it adds a layer to a task the user already does competently. Axis 2: the output has to be correct, or plausible is enough. The shippable cell is "replaces avoided task" where "plausible is enough." Every other cell is a demo with a roadmap ticket attached.
Why This Applies to 17 AI Features on a Roadmap
Meta's internal token-consumption leaderboard was gamed immediately. Engineers wrote scripts that burned millions of tokens doing nothing. Meta shut it down. The same pattern shows up in product dashboards. Teams count tokens consumed and sessions opened and call it adoption. The durable numbers are harder to collect: retention of AI-assisted workflows after 30 days, time-to-first-useful-output, percentage of AI output that ships to production without rewrite.
Multiple sources confirm the bottleneck has moved from engineering capacity to discovery and specification quality. When shipping takes an afternoon, shipping the wrong thing takes an afternoon too. Feature count goes up. Time-to-value does not.
The Margin Forcing Function
Microsoft gets OpenAI's technology at preferential rates and still admitted Copilot inference costs drag margins. Every AI feature carries an ongoing compute cost that scales with usage, not with value delivered. A feature with 5% engagement and 100% inference cost on every page load is burning money on 95% of impressions. Microsoft is not cutting features because users hate them. It is cutting features because each impression has a marginal cost and most impressions don't earn it back.
The Organizational Response
Nadella merged consumer and enterprise Copilot under a single EVP, Jacob Andreou. The diagnosis is governance, not product. When every product team independently bolted a chatbot onto its surface, nobody owned the holistic experience or the total inference bill. Teams with distributed AI feature ownership and no central quality gate are on the trajectory Microsoft was on six months ago.
What to do
Map every AI feature in your product onto the 2x2 grid (replaces-avoided-task vs. adds-layer, and plausible-output vs. must-be-correct) by end of this sprint
Pull 30-day retention and usage depth for each AI feature shipped in the last 90 days — replace token/session metrics on leadership dashboards
Propose a centralized AI product owner or quality council to leadership using Microsoft's 81-product cautionary tale
Kill or pause at least 3 AI features that show no usage lift after 4+ weeks in production