Product & Strategy

The Product Desk

The Signal

Microsoft pulled Copilot from five Windows 11 apps after 'near-universal' backlash

The 'add AI everywhere' playbook is being punished from every direction simultaneously. If your AI roadmap is still framed around 'time saved,' NVIDIA's Shraddha Sridhar just showed you the ceiling: 30%. The teams redesigning entire workflows around AI — with traceability as a P0, not a v2 — are the ones pulling away.

In Play

  1. AI Integration Backlash Reaches Critical Mass

    Microsoft retreated from Copilot in 5 Windows apps. Xbox leads with 'no AI slop.' Hachette pulled a book on AI suspicion alone. Gamers revolted against DLSS. NVIDIA's internal AI failed without traceability. Consumer and professional hostility to shallow AI is now a first-order product constraint — not a PR issue.

    Ask Clarity
  2. Inference Demand Explodes 1,000,000x — Token Budgets Are Wrong by Orders of Magnitude

    Per-user token consumption jumped 1,000x in 2 years (100K→100M/day), peaking at 870M in one day via multi-agent architectures. Aggregate inference demand expanded ~1,000,000x. A 44 GW data center power shortfall persists through 2028, but Vera Rubin + Groq promises 35x throughput/watt in H2 2026.

    Ask Clarity
  3. AI Benchmark Credibility Crisis — Half of 'Passing' Code Won't Ship

    METR found ~50% of ~300 AI PRs passing SWE-bench Verified wouldn't actually be merged — failures in code quality, broken surrounding code, and missed functionality. Meanwhile, the 'almost perfect' supervision paradox means high AI accuracy trains humans to stop checking. Your PRD accuracy claims likely have a ~50% overstatement.

    Ask Clarity
  4. AI Monetization Reckoning — Vague Strategy Now Destroys Value

    Alibaba and Tencent lost $66B in 24 hours after earnings showed heavy AI spending with no monetization path. Enterprise AI pricing ($200/seat/mo) is 10x consumer — but only if you can quantify ROI. OpenAI is pivoting to enterprise because consumer AI economics aren't working. VC money is pulling back from consumer AI plays.

    Ask Clarity
  5. Federal AI Framework Proposes Single National Standard

    The White House released an AI framework that would override state laws, limit platform liability, shift child safety to parents, and give startups broad protections. If enacted, teams building for California/Illinois/EU patchwork compliance may reclaim significant engineering bandwidth. Still a proposal — not legislation.

    Ask Clarity

Deep Dives

The AI Integration Backlash Is Now a Product Constraint — And NVIDIA's Internal Failure Shows the Way Out

This week, the evidence became undeniable: shallow AI integration is being systematically rejected — by consumers, by professionals, and by the internal engineering teams at the company most invested in AI's success. The question for PMs is no longer whether to add AI, but how to add it without triggering the backlash that just forced Microsoft into the most public product retreat of the year.

The Backlash Data Is Now Overwhelming

Microsoft pulled Copilot entry points from Snipping Tool, Photos, Widgets, and Notepad after what they acknowledged as 'near-universal negative user feedback.' The replacement features — a movable taskbar, fewer forced restarts, faster File Explorer — are almost embarrassingly basic. Microsoft neglected core UX hygiene while chasing AI integration, and users noticed. In the same week, Xbox's new leader Asha Sharma (handpicked by Nadella) made her first public promise: 'No Soulless AI Slop.' When the company that bet $13B on OpenAI is marketing against AI, the positioning landscape has fundamentally shifted.

The creative economy is reacting even more viscerally. Hachette pulled a published novel from stores on mere suspicion of AI involvement — without proof. Conan O'Brien mocked AI at the Oscars. Gamers revolted against NVIDIA's DLSS update, calling it 'the same boring Instagram filter,' and Jensen Huang's response — telling users they were 'completely wrong' — made it worse.

NVIDIA's Internal Proof: Even Engineers Won't Use AI Without Traceability

The most instructive data point didn't come from consumers — it came from inside NVIDIA. Their chip-design team tried a fine-tuned AI domain expert in 2023. It failed completely. Not because the model was bad, but because hardware engineers demanded traceability — the ability to trace every AI output to a source document. Product Lead Shraddha Sridhar rebuilt the system around curated documents, source attribution, and verifiability. Only then did adoption take off.

"We fixed the problem of traceability and verifiability, which meant engineers would trust their responses. And that was key to driving adoption." — Shraddha Sridhar, NVIDIA

The 30% Ceiling and the Three-Tier Framework

Sridhar outlined three AI deployment tiers that should reshape how you measure success:

  1. Individual productivity — your copilot sidebar. Caps at ~30% time saved.
  2. Team-level scaling — shared AI workflows that multiply team output.
  3. Capability expansion — AI enables things that were previously impossible.

Most product teams are entirely in Tier 1. The electric motor analogy crystallizes why that's insufficient: motors arrived in factories in the 1880s, but productivity gains didn't materialize until the 1920s — because early adopters just swapped the power source and kept the old floor plan. Real gains required redesigning the factory. If you're adding an AI chat panel to your existing UI, you're in the 1880s. The compression in software is faster (3-5 years, not 40), but the principle is identical.

The Synthesis No Single Source Provides

Cross-referencing the consumer backlash with NVIDIA's internal experience reveals a unified pattern: AI that doesn't serve a specific, traceable user need gets rejected by every audience. Consumers reject it as 'slop.' Engineers reject it as untrustworthy. Markets reject it as unmonetizable (see: Alibaba/Tencent's $66B wipeout). The path forward isn't less AI — it's redesigned AI that treats traceability as table stakes, measures capability expansion rather than time saved, and integrates invisibly where it should be invisible.

What to do

  1. Audit every AI feature in your product for traceability — can users trace each output to source data? Add source attribution as P0 to current sprint for any that lack it.

  2. Score every consumer-facing AI feature on a 'slop risk' rubric: Does it homogenize output? Can users opt out? Is AI labeled or invisible? Does it optimize for the metric users actually care about? Complete by end of sprint.

  3. Reframe your top 3 AI feature success metrics from 'time saved' to 'capabilities unlocked' and present the 3-tier framework to leadership this quarter.

  4. Scope a v2 AI-native architecture for one core workflow — redesigned assuming AI is a first-class resource, not retrofitted onto existing UX.

Token Demand Hit 1,000,000x — Your Consumption Model Is Off by Three Orders of Magnitude

Forget the cost-per-token improvements we covered Saturday. The bigger story is on the demand side: per-user token consumption grew 1,000x in under two years, aggregate inference demand expanded roughly 1,000,000x, and the infrastructure to serve it faces a 44 GW power shortfall through 2028. If you're modeling AI costs as a fixed line item, you're building on sand.

The Consumption Explosion Has Hard Numbers

Azeem Azhar's documented personal usage went from ~100K-150K tokens/day in mid-2024 to 100M tokens/day in March 2026 — a 1,000x increase. On a single heavy Monday, his multi-agent system (one chief-of-staff agent orchestrating four specialized sub-agents for research, portfolio management, editorial, and frameworks) consumed 870 million tokens in a single day. This isn't theoretical. It's one power user with a four-agent setup.

Most large companies treat token budgets as an IT cost center. Azhar argues this is 'dangerously behind the curve' — tokens should be treated as fundamental as electricity or office space.

The aggregate math is even more dramatic: 10,000x more compute per interaction (as users shift from chat to reasoning models and agentic systems) multiplied by 100x more users deploying at scale equals a million-fold expansion in inference demand over roughly two years.

Infrastructure Can't Keep Up — But Relief Is Coming

GPUs are architecturally mismatched for inference workloads. During decode, thousands of cores sit idle waiting on memory bandwidth. This is why NVIDIA valued Groq (inference-specialized chips) at $20 billion. The combined Vera Rubin + Groq architecture, due H2 2026, promises 35x throughput per megawatt versus current Blackwell. Meanwhile, NVIDIA has locked up 70% of TSMC's 3nm capacity, ASML can only produce ~700 EUV machines per year, and Morgan Stanley projects the 44 GW data center power shortfall persists through 2028.

The Pricing Paradox for PMs

Jensen Huang stated he'd be 'deeply alarmed' if a $500K developer spent less than $250K on AI tokens annually. That's NVIDIA talking its book, but directionally it signals where the market is heading: AI compute costs will simultaneously drop per unit and increase in total spend as usage expands. Your financial model needs two curves, not one: a declining cost-per-token curve and an exponentially rising consumption curve. The intersection determines whether your AI features are profitable.

A new demand paging technique for LLMs (reducing memory by 90% within 1% accuracy) further confirms the cost curve is bending. Features you marked 'too expensive to run at scale' six months ago need re-evaluation against both the 35x hardware improvement and these software optimizations.

What This Means for Enterprise Positioning

If tokens are a cost, every AI feature you ship increases perceived expense. If tokens are a productive input, every AI feature increases capacity. The framing determines whether your champion gets a bigger budget or gets audited. The first product in each enterprise category that successfully makes the productivity-input argument captures the procurement conversation for years.

What to do

  1. Model agentic consumption scenarios at 100x-1,000x your current per-user token assumptions — identify at what multiple your unit economics break. Complete by next planning cycle.

  2. Reprioritize AI features previously shelved for cost using 35x throughput improvement as a planning assumption for H2 2026+. Build a declining cost curve into your models.

  3. Reframe enterprise AI positioning from 'infrastructure cost' to 'productivity input' in sales enablement materials and pricing pages.

  4. Prototype a multi-agent architecture for your highest-value workflow to measure real-world token consumption patterns in your domain.

Half Your AI's 'Passing' Code Won't Actually Ship — The Benchmark Credibility Crisis Hits Your PRDs

If you've cited a benchmark score in a PRD, a board deck, or a vendor evaluation in the last six months, this week's METR research just undermined your numbers. And the implications extend far beyond code generation into every AI feature where humans nominally supervise high-accuracy outputs.

The METR Finding: 50% Overstatement

METR evaluated roughly 300 AI-generated pull requests that passed SWE-bench Verified's automated grader. The result: approximately half would NOT actually be merged by real repository maintainers. Failures included code quality issues, broken surrounding code, and core functionality that the test suite simply missed. This isn't a minor calibration error — it's a ~50% overstatement baked into every accuracy claim built on benchmark scores.

The implications are immediate:

  • Any vendor selling you on SWE-bench scores needs to show real-world merge rates, not benchmark pass rates
  • Your internal AI features need human-review validation layers before you can credibly claim accuracy
  • Stripe's principle that 'tool curation matters more than tool quantity' looks like hard-won wisdom, not opinion

The 'Almost Perfect' Supervision Paradox

This benchmark gap compounds with a more fundamental design problem articulated by Raffi Krikorian (formerly head of Uber's self-driving unit), who crashed his Tesla and wrote what may be the year's most important UX insight:

"We are asking humans to supervise systems designed to make supervision feel pointless. A machine that works almost perfectly? That's where the danger lies."

Every PM shipping AI copilot features, AI-assisted moderation, or AI-generated content with human review needs to internalize this. Your highest-performing AI features may be your most dangerous — because they've trained users to stop paying attention. The failure mode isn't AI getting worse; it's AI being good enough that humans stop catching the 50% that shouldn't ship.

A Potential Fix: Autoresearch Self-Improvement Loops

One bright spot: a new 'autoresearch' pattern where AI agents iteratively optimize their own outputs using binary yes/no checklists improved landing page copy from 56% to 92% pass rate in 4 rounds and page load from 1100ms to 67ms, with zero human intervention. For any AI feature with measurable quality rubrics, this self-improvement loop is worth experimenting with — it addresses the benchmark gap by grounding AI in your specific quality criteria rather than generic test suites.


Cross-Source Tension Worth Flagging

There's a contradiction in this week's signals: companies like Stripe, Ramp, and Coinbase are deploying autonomous coding agents that pick up tickets and open PRs, while METR proves half of AI code that 'passes' won't ship. The reconciliation is in curation, not capability: Stripe runs ~15 curated tools with AGENTS.md files encoding org-specific conventions. They've solved the quality problem by constraining the agent, not by trusting the benchmark. Your agent deployment strategy needs the same rigor.

What to do

  1. Replace benchmark-based accuracy claims with real-world evaluation metrics in all active PRDs and vendor assessments. Flag any claims citing SWE-bench for mandatory revision this sprint.

  2. Audit your product for anywhere AI accuracy exceeds 95% but humans nominally supervise — redesign the supervision UX to re-engage human attention at critical decision points.

  3. Pilot an autoresearch-style self-improvement loop on one AI feature with measurable quality rubrics (content generation, search, recommendations) this quarter.

  4. If deploying coding agents, implement AGENTS.md-style convention files encoding your team's quality standards and architectural decisions into every agent run.

The bottom line

The 'add AI everywhere' era ended this week from both directions: consumers systematically reject it (Microsoft retreated from five apps, Xbox banned 'AI slop,' Hachette pulled a book on suspicion alone), markets punish it ($66B evaporated from Alibaba and Tencent for vague AI strategies), and METR proved half of AI-generated code that 'passes' benchmarks wouldn't actually ship. Meanwhile, per-user token demand hit 1,000x in under two years — meaning even the teams that get AI right will need to completely remodel their cost assumptions. The PMs who win from here stop measuring 'time saved,' start measuring 'capabilities unlocked,' add traceability to every AI output before anything else, and build their financial models around a consumption curve that is three orders of magnitude higher than what's in their current spreadsheets.