Product & Strategy

The Product Desk

The Signal

Stripe's Machine Payments Protocol went live this week: 894 AI agents executed 31

Meanwhile, Databricks data from 20,000+ orgs proves companies with AI governance frameworks push 12x more projects to production. The two signals converge: your product needs to be both discoverable by agents and governed enough to ship AI features at pace.

In Play

  1. Agent-Native Commerce Goes Live on Stripe

    Stripe and Tempo shipped MPP in March 2026 — 894 agents, 31K+ transactions, 60+ services in week one. Payment is embedded in the HTTP request itself; no accounts, no API keys, no checkout. Visa shipped a CLI tool for agent payments. The 'headless merchant' archetype is live and funded.

    Ask Clarity
  2. Enterprise AI Governance = 12x Production Multiplier

    Databricks telemetry (20K+ orgs, 60%+ Fortune 500) shows governance frameworks correlate with 12x more AI projects reaching production. Separately, 29% of Fortune 500 are now paying AI startup customers — but coding dominates by 10x, and Harvey hit $200M ARR with sub-50% model accuracy. Revenue tracks workflow fit, not model capability.

    Ask Clarity
  3. Open-Source Coding Model Tops Proprietary — First Time Ever

    Z AI's GLM-5.1 scored 58.4 on SWE-Bench Pro, dethroning GPT-5.4 and Opus 4.6 — the first open-source model to claim #1 on this benchmark. It's MIT-licensed, runs 8-hour autonomous sessions with 1,700 tool calls, and was trained entirely on Huawei Ascend chips with zero Nvidia silicon. Your API cost assumptions may be 3x too high.

    Ask Clarity
  4. Development Cadence Collapse: Shape Up's Creator Killed It

    DHH — who wrote Shape Up's 2-month cycle methodology — now calls it obsolete due to AI-accelerated velocity. Separately, one non-coder shipped 70K LOC in 7 weeks with 85% test coverage and 9.5/10 code health. But research on 200 programmers shows AI assistants reduce persistence through hard problems by 25%. Speed is real; depth is at risk.

    Ask Clarity
  5. Post-Quantum Deadline Moves Up 6 Years

    Cloudflare set 2029 as its post-quantum cryptography deadline — a 6+ year compression from previous 2035+ estimates. Google revealed a breakthrough algorithm for elliptic curve crypto, and Oratomic showed P-256 crackable with just 10K qubits. If your product uses TLS, JWT, or encrypted data at rest, scoping starts now.

    Ask Clarity

Deep Dives

Agent-Native Commerce Is Live — Your Subscription Model Has a Per-Request Competitor

The Headless Merchant Archetype Is Real and Transacting

Stripe and Tempo co-built the Machine Payments Protocol (MPP), which went live in March 2026. In its first week: 894 AI agents executed 31,000+ transactions across 60+ API-only services, with per-request pricing from $0.003 to $35. No human-facing UI. No user accounts. No API keys. Payment is embedded directly in the HTTP request — the transaction is the authentication.

When payment is the authentication, there's nothing to lock in. The stickiest part of SaaS — the account — disappears.

The services range from SEC filing search to image generation (fal.ai offers 600+ models at fractions of a cent) to physical letter mailing. Visa released a CLI tool for agent payments alongside MPP, and the protocol supports cards, stablecoins, and Lightning in a single flow. When Stripe co-builds a protocol and Visa ships developer tooling in the same quarter, the rails aren't the bottleneck anymore.


Why This Threatens Subscription SaaS Directly

Consider the unit economics inversion. If you sell an image generation subscription at $10/month, an AI agent doesn't need your subscription — it needs one image right now, and fal.ai delivers it at $0.003 with zero friction. The agent arrives with intent fully formed — it knows what it needs, what format, what it'll pay. Your brand doesn't matter. Your onboarding flow doesn't matter. What matters: can the agent read your schema, call your endpoint, and get a result in one HTTP round trip?

This dynamic is amplified by a structural finding across multiple analyses: AI agents systematically prefer open and free software over closed commercial alternatives. When agents make tool-selection decisions, they reach for services they can access without human-gated signups, license keys, or sales calls. Your beautifully-gated enterprise product is invisible to the fastest-growing class of buyers.

The Unsolved Problem: Agent Discovery

Today it's a 60-service directory. In a year, it could be 60,000. There's no SEO, no app store, no search equivalent for agent-consumable services yet. Whoever builds 'Google for agent commerce' captures the most valuable chokepoint in this stack. The discovery problem is either your biggest threat or your biggest opportunity.

The Risk Side

Micropayment unit economics at $0.003/request require staggering volume — 31K transactions in week one is roughly $93 in revenue at the floor price. Regulatory exposure for autonomous stablecoin transactions without KYC will attract scrutiny. Protocol fragmentation (MPP vs. x402 vs. Visa) could slow adoption. But Stripe doesn't co-build protocols for concepts that don't scale, and Visa doesn't ship CLI tools as experiments.

What to do

  1. Model a headless competitor scenario this sprint: what happens if someone offers your core API capability at per-request pricing with no signup?

  2. Ship a machine-readable service schema (pricing, capabilities, I/O formats) alongside your existing API docs by end of Q2

  3. Evaluate adding a per-request pricing tier alongside your subscription model — start with your lowest-friction API endpoint

  4. Brief your legal team on autonomous agent transaction implications, especially for regulated data

Governance Is the 12x Velocity Lever — And Enterprise AI Adoption Data Just Got Definitive

The Production Deployment Gap Has a Number

Databricks' 2026 State of AI Agents report — covering 20,000+ organizations including 60%+ of the Fortune 500 — drops the most consequential enterprise AI finding this quarter: companies with AI governance frameworks push 12x more projects to production than those without. The mechanism is intuitive: teams with evaluation pipelines, rollback mechanisms, and approval workflows can confidently ship. Teams without get stuck in POC purgatory and executive risk aversion.

Governance isn't your speed bump — it's the thing that lets you ship 12x more stuff.

This converges with a16z's enterprise adoption data published April 8: 29% of Fortune 500 and ~19% of Global 2000 are now live, paying customers of AI startups — clearing the highest bar (top-down contracts, successful pilot conversion, production deployment). For context, cloud computing took 7–8 years to reach comparable penetration. AI did it in 3.3 years.


The Use-Case Hierarchy Changes Your Prioritization

Coding dominates enterprise AI deployment by nearly 10x over support and search. a16z's 5-trait framework explains why:

  1. Text-based workflows
  2. Rote/repetitive tasks
  3. Natural human-in-the-loop
  4. Limited regulation
  5. Clearly verifiable outputs

Before building any AI feature, score it against these five traits. Anything below 3/5 should be deprioritized. Coding scores 5/5 — the code compiles and passes tests, or it doesn't. That's why Cursor, Claude Code, and Codex are outpacing projections.

The Capability-Revenue Gap Is Your Whitespace Map

Harvey hit ~$200M ARR in legal AI while models score below 50% against human lawyers on GDPval. You don't need best-in-class model performance — you need a copilot that makes expensive humans more productive. The inverse is your opportunity: accounting/auditing just jumped ~20% on GDPval in 4 months, and police/detective work improved ~30%, yet neither has a breakout AI company. The capability gap is closing fast; the revenue gap is wide open.

The Kill-or-Prove-It Moment

Multiple enterprise sources confirm the shift from 'deploy AI because the board said so' to 'prove AI works or lose your budget.' MassMutual and Mass General Brigham both moved to production by centralizing governance, enforcing success metrics per use case, and actively killing low-value pilots. If your customer's governance team can't see your feature's impact in their dashboard, you won't survive their next portfolio review. 21% more teams report AI cost savings vs. 2024, and 91% of service management orgs say AI saves money — but features without finance-verifiable outcomes are getting killed faster than in 2024.

What to do

  1. Add an AI governance workstream to your current roadmap: evaluation frameworks, human-in-the-loop approval flows, model monitoring. Use the 12x stat as your business case in the next planning review.

  2. Score every AI feature on your roadmap against the 5-trait adoption framework. Kill or deprioritize anything below 3/5.

  3. Map the GDPval capability-vs-revenue gap for your vertical and bring the analysis to your next strategy review

  4. Build an AI outcomes dashboard that surfaces finance-verifiable metrics (cost savings, MTTR reduction, resolution time) to your enterprise buyer's governance team

Your Sprint Cadence Is Probably Wrong — Three Independent Proof Points This Week

Shape Up's Creator Killed Shape Up

DHH — who created the Shape Up methodology at 37signals and wrote its definitive book — told Lex Fridman that 2-month development cycles are now too slow. He went from typing every line by hand (October 2025) to barely writing code at all (April 2026), running Gemini 2.5 and Opus 4.5 in parallel, reviewing diffs via Lazygit. He describes the shift as 'wearing a mech suit.' When your methodology's own creator calls it obsolete due to AI, it's not an anecdote — it's a leading indicator for your planning cadence.

37signals operates with 10 designers and 20 engineers (1:2 ratio), where designers own product definition, UX, and build the first version. DHH believes AI tools are pushing the industry toward this designer-builder model. For PMs: tactical roles that primarily write tickets and manage backlogs lose their reason to exist when a designer with Claude Code goes from insight to working prototype in a day.


70K LOC, 7 Weeks, One Person, Zero Hand-Written Code

Luca Rossi's Tolaria project produced a production-grade 70,000-line codebase with 3,000 tests, 85% coverage, a 9.5/10 CodeScene health score, and 40+ Architecture Decision Records — all in ~7 weeks. He operates as 'CEO' of AI agents: OpenClaw produces specs, Claude Code executes, CI gates on code health and test coverage serve as the primary quality controls. Not code review. Not human QA.

At 30 commits per day, no human can meaningfully review each change. Automated guardrails catch regressions — the new quality primitive is CI gates, not code review.

The codebase grew 2.5x in a single month (20K to 70K LOC) while code health improved. This isn't 'move fast and break things' — it's 'move impossibly fast and things don't break.'

The Persistence Problem: Speed Without Depth

A counterweight emerged from research on 200 programmers: AI coding assistants make developers 25% less likely to persist through challenges. Amazon has already codified this into policy, restricting junior devs from shipping agent-generated code without senior review. DHH confirms that senior engineers benefit 'a lot more' from AI than juniors at 37signals. The implication: sprint velocity may look great in aggregate while hiding that complex architectural work stalls as developers bail to AI-generated workarounds.

The CLI-as-Agent-Interface Insight

DHH is building CLIs for every 37signals product because CLIs let agents chain tools together — an agent checks Sentry, writes a fix, posts a PR to GitHub, and reports to Basecamp, all via CLI. He calls CLIs 'the ultimate AI interface.' Products that are agent-chainable get embedded into automated workflows and become stickier. Products that require a GUI for core actions become invisible to the agent layer.

What to do

  1. Audit your current sprint/cycle duration against actual throughput this sprint — if teams consistently finish early, propose a compressed cadence experiment (1-week cycles or continuous delivery with weekly prioritization)

  2. Run a controlled build experiment: assign one senior PM or tech lead to attempt building your next internal tool solo with AI coding tools in a 2-week sprint, with CI gates on coverage and code health

  3. Establish AI code quality gates by seniority level — juniors require senior review for agent-generated code, seniors require CI automation (coverage thresholds, health scoring)

  4. Scope a CLI or agent-accessible API surface for your product's core workflows by end of Q3

The bottom line

Agent-native commerce went live on Stripe this week — 894 AI agents, 31,000 transactions, $0.003/request, zero signups — and Databricks proved governance (not features) is the 12x multiplier for getting AI to production. Meanwhile, an open-source model just topped every proprietary competitor on coding benchmarks for the first time, and Shape Up's own creator declared his 2-month cycles obsolete. The winning PM this quarter invests a sprint in governance infrastructure before writing another AI feature, ships a machine-readable service schema so agents can find you, and runs their own build experiment before someone else proves their team is 3x too big.