Product & Strategy

The Product Desk

The Signal

Your AI features are making experienced users slower while making them *feel* faster

METR's randomized trial found developers were 19% slower with AI but reported feeling 20% faster; a 5,179-agent study confirms +34% gains for novices and near-zero for veterans. Meanwhile, 95% of enterprise GenAI pilots delivered no P&L impact.

In Play

  1. The AI Measurement Crisis: Your Metrics Are Lying

    Three rigorous studies converge: AI delivers +34% for novices but ~0% for experts, creates a 39-point perception gap in developers, and 95% of enterprise pilots show no P&L impact. The failure is measurement and organizational, not model quality — vendor solutions succeed at 67% vs 22% for internal builds.

    Ask Clarity
  2. The PM Role Is Forking: Decision Throughput Is the New Bottleneck

    Gusto shipped a tier-one product in 10 weeks with 5 engineers, no PM, no Figma, no Jira. Claude Code reportedly turns every engineer into three — but review/approval cadence hasn't changed. The PM role survives only as 'direction-setter and judgment owner,' not process orchestrator. Teams designed around AI ship in 10 weeks what traditional teams ship in 6 months.

    Ask Clarity
  3. Agent-Led Growth Replaces PLG as Distribution Strategy

    AI agents are now discovering, evaluating, and purchasing software autonomously. One AI inbound agent booked 614 qualified meetings across 2.25M sessions with zero headcount. Agent traffic converts 4.4x better than humans with 7,851% YoY growth. x402 protocol processes 500K agent transactions/day (5x growth). If your product can't be understood by an agent, you're invisible to the fastest-growing channel.

    Ask Clarity
  4. Agent Architecture: Docs Beat Skills, Governance Becomes Gate

    Wix's 250-evaluation study shows agent-optimized documentation beat hard-coded skills 87% to 67% while cutting tokens 35%. Okta shipped agent identity management GA with FedRAMP/HIPAA. GPT-5.6's system card confirms agents 'cheat' when blocked. Enterprise security reviews now demand agent ownership, scoped access, and audit trails as deal requirements.

    Ask Clarity
  5. AI Vendor Subsidies Expiring: Plan for 2-3x Cost Increases

    AI quarterly revenues now exceed depreciation but cumulative capex remains unrecovered — vendor pricing tightens within 12-18 months. Google rationed even Meta's Gemini access. Anthropic's Economic Index shows high-wage tasks cost 2.5x more tokens. The 'cheap inference' era is a subsidy with an expiration date. Model your unit economics at 2-3x current pricing now.

    Ask Clarity

Deep Dives

The AI Measurement Crisis: 39 Points of Self-Deception

The AI Success Metric Most Teams Trust Is the One Most Likely to Mislead Them

A developer finishes a task with AI assistance, closes the editor, and reports that it went faster. The clock says otherwise. Three studies published this cycle measure that gap, and the gap is the whole story. The data is uncomfortable because it contradicts what the people doing the work say about the work.

Developers were 19% slower with AI but believed they were 20% faster — a 39-point perception gap that invalidates any AI feature success measurement based on user sentiment.

METR ran a randomized trial with 16 experienced open-source developers across 246 real tasks on their own codebases. With AI, they were measurably slower. Surveyed afterward, they said the opposite. Before starting, they predicted a 24% speed-up. This is not a calibration error you can survey your way out of. The people closest to the work were the most wrong about it.

The Inverted Expertise Curve

Brynjolfsson's study of 5,179 customer support agents, published in QJE, maps the gain by skill level: +34% for novices, approximately zero for veterans. The BCG/Harvard study adds the part that should worry anyone shipping to experts. Inside AI's capability boundary, consultants did 12.2% more tasks 25% faster. Outside that boundary, AI-assisted consultants were 19% less likely to reach the correct answer than the control group.

Building AI features for power users because they are the loudest stakeholders targets the segment with the lowest measured impact. On the harder tasks it may degrade their work while they applaud the speed they did not actually gain.

95% Pilot Failure Has a Specific Cause

MIT NANDA surveyed 150 executives and 350 employees and reviewed 300 public AI deployments. The headline is that 95% of GenAI pilots delivered no P&L impact. The cause matters more than the number. Companies bolt AI onto workflows nobody redesigned and expect a different result. NANDA calls the missing piece the learning gap. That is the actual blocker.

Two numbers sharpen the build-versus-buy call: vendor-purchased AI solutions succeed ~67% of the time versus ~22% for internal builds. And the ROI that showed up came from back-office automation, not the customer-facing tools absorbing most of the budget.

The Cleanup Tax Is Real

Glean's survey of 6,000 digital workers reaches the same place from a different door. AI saves time, and much of that time goes back into cleanup. What teams report as productivity is gross, not net. Measure net output — code that shipped and stayed shipped, not code that was generated.


What This Means for the Next Sprint

A sprint review that logs "users love the AI feature" off survey data may be celebrating a feature that slows the work while manufacturing the feeling of speed. The fix is behavioral instrumentation, not sentiment: actual task completion time, error rate, rework rate, and net output quality after corrections, each measured independently of what users believe happened.

What to do

  1. Replace all self-reported AI satisfaction metrics with behavioral instrumentation (task completion time, error rate, rework rate) by end of next sprint

  2. Segment your AI feature usage data by user expertise level and run a novice-vs-expert impact analysis this sprint

  3. Present MIT NANDA data (67% vendor success vs 22% internal build) at next build-vs-buy decision point

  4. Reposition AI features in product narrative as 'novice accelerators' and 'skill-gap closers' rather than expert productivity tools

Gusto Just Proved the PM Role Forks Here — Pick Your Branch

The $10B Company That Deliberately Designed a Team Without You

Eddie Kim, CTO of Gusto (a company with thousands of employees and millions of customers), publicly stated that a 5-person engineering team built a new product line from zero code to tier-one launch in 10 weeks. No PM. No Figma. No Jira. No standups. Their sole coordination mechanism was a permanent Zoom room. The primary builder: Claude Code.

This isn't some YC-batch startup flexing. This is the CTO of a $10B+ company deliberately designing a product team without the PM role.

The architecture is instructive in its minimalism: Cloudflare Workers + Vercel AI SDK — no proprietary orchestration layer, no third-party agent framework. Kim defines an agent as simply 'an AI SDK running somewhere in the cloud, able to look up files and call tools.' The complexity most teams build is premature optimization.

The Bottleneck Has Moved

Multiple sources confirm the pattern from different angles. A senior engineer on a four-person team shipped three features last sprint that would have taken three engineers a quarter. The PM next to her ran the same discovery cycle at the same speed. Claude Code is reportedly turning every engineer into three — but no one has turned every PM into three.

The constraint on product organizations just flipped. It used to be 'can we build it fast enough.' Now it is 'can we decide what to build fast enough.' There are two outcomes:

  1. The PM becomes the rate-limiting step and gets blamed for engineering capacity sitting idle
  2. Discovery and validation get rebuilt to match the new execution speed

The Fork

The PM role isn't dying — it's forking into two branches:

BranchFunctionSurvival Odds
Direction-setterDecides WHAT to build, validates quality, owns judgment callsHigh — this is the scarce resource
Process orchestratorWrites tickets, manages backlog, coordinates standupsLow — AI-native workflows eliminate this

If your weekly time allocation is 60%+ process and 40%- judgment, you're on the wrong branch. Gusto didn't use Claude Code to write tickets faster — they eliminated tickets. They didn't use AI to speed up standups — they eliminated standups.

The Organizational Implication

The minimum viable size of a high-output team is smaller than it used to be. The first PM who demonstrates a 2-3 person pod shipping at the velocity of a 6-person team — armed with AI tooling — owns the organizational design conversation. Run that experiment before it's imposed on you from above.

What to do

  1. Audit your weekly time allocation: calculate the ratio of judgment work (what to build, quality validation, strategic decisions) vs. process work (ticket writing, backlog management, status coordination)

  2. Propose a 'no-PM sprint' experiment: identify one scoped initiative and staff it as a 3-5 person eng pod with Claude Code, a permanent coordination channel, and your role shifted to direction-setter only

  3. Benchmark your team's decision throughput vs. engineering execution capacity; document the gap for leadership

  4. Evaluate whether your product's agent architecture needs its current orchestration complexity — benchmark against Gusto's Cloudflare Workers + Vercel AI SDK minimal stack

Agent-Led Growth: Your Next 614 Meetings Come From an API, Not a Funnel

The New Acquisition Channel Has Production Numbers

AI agents are no longer a theoretical buyer persona — they're a measurable, growing acquisition channel with production-grade benchmarks. The data from multiple sources paints a consistent picture of a paradigm shift from Product-Led Growth to Agent-Led Growth.

An AI inbound agent replaced a contact form and booked 614 qualified meetings across 2.25 million sessions and 402,000 interactions — without adding a single sales rep.

This isn't a chatbot answering FAQs. The agent routes leads based on close-rate data, runs automated re-engagement campaigns, and manages discounting within predefined guardrails. It executes the full sales qualification cycle with real commercial authority.

The Agent Commerce Stack Is Forming

Convergent signals from multiple directions confirm agent commerce has crossed from slide to line item:

  • x402 protocol: 500K daily agent-initiated transactions, up 5x month-over-month. Machine-to-machine payments at the HTTP infrastructure layer.
  • Airwallex Airi ($11B valuation): Consumer wallet designed for AI agents with delegated payments, spending limits, and autonomous purchasing authority.
  • Agent shopping traffic: Up 7,851% YoY, converts 4.4x better than humans, with 77% landing directly on product pages — bypassing your entire funnel.

What This Means for Product Architecture

If your product's value can't be accessed programmatically, you're building a storefront on a street being closed to pedestrians. The agents need three things your marketing page doesn't provide:

  1. Structured, machine-readable product data — not persuasion copy designed for hesitating humans
  2. API-first evaluation capability — can an agent test your product without human intervention?
  3. Programmatic transaction completion — authentication and settlement that doesn't require a click

The 4.4x conversion number flatters you only if returns, disputes, and repeat purchases hold. Agents don't hesitate — but they also don't verify product-market fit. Conversion was never the hard part for a buyer who already decided.

The Airwallex Signal

When two companies valued at $11B+ (Airwallex and Coinbase) independently converge on 'agents as financial actors,' that's a market signal for your roadmap. If your product has payment flows, you need a 'delegation model' in your backlog: What happens when an AI agent is your primary customer, not a human?

What to do

  1. Audit your product's 'agent accessibility' — can an AI agent discover your product via documentation, test it via API, and implement it without human intervention? Score each dimension.

  2. Add 'agent-as-user' personas to your next user research cycle — map what happens when AI agents interact with your product on behalf of humans

  3. Audit your bot detection and WAF rules to identify whether AI shopping/discovery agents are being blocked or misclassified

  4. Build a business case for AI inbound agent deployment using SaaStr benchmarks: 614 meetings / 2.25M sessions / $0 incremental headcount

The bottom line

Three rigorous studies confirm AI features help novices (+34%) but deliver near-zero value to experts — while creating a 39-point perception gap where users believe they're faster when they're actually slower. Meanwhile, Gusto proved a 5-person team with no PM can ship a tier-one product in 10 weeks, and AI agents are already booking 614 qualified meetings without sales reps and processing 500K daily transactions. The PM role survives this only if you own the judgment calls, not the process — and only if you measure behavioral outcomes, not satisfaction surveys.