Product & Strategy

The Product Desk

The Signal

Your engineering team's AI toolchain flipped overnight

Meanwhile, OpenAI is building a GitHub competitor it plans to sell commercially. If you haven't recalibrated your roadmap capacity estimates and platform dependencies against these numbers, your sprint velocity baselines and integration strategy are already stale.

In Play

  1. AI Coding Tools Reshape Engineering Capacity and Platform Dependencies

    A 906-person senior engineer survey confirms AI coding tools have crossed from augmentation to operating model (95% weekly use, 56% doing 70%+ of work with AI), Claude Code displaced GitHub Copilot in 8 months, Cursor hit $2B ARR, and OpenAI is building a GitHub competitor — collectively forcing immediate recalibration of engineering capacity planning, platform integration strategy, and build-vs-buy decisions.

    Ask Clarity
  2. SaaS Pricing Models Breaking Under AI Commoditization

    AI has collapsed product build costs from ~$200K to as low as $20, mid-tier models match flagships at 40% less cost, and only 5.5% of ChatGPT's 900M users pay — converging to make seat-based pricing, feature-parity moats, and capability-driven marketing obsolete, with competitive defensibility migrating to proprietary data, workflow depth, and emotional design.

    Ask Clarity
  3. AI Agent Infrastructure Crystallizing Into a Real Stack

    Agent infrastructure is shipping across payments (agentcard.sh), identity (actors.dev), hosting (iocaihost), stateful runtimes (OpenAI on AWS), and GTM automation (a16z founders running 5-agent sales orgs via iMessage) — forming a distinct stack that demands your product design for agent-as-user, not just AI-as-feature.

    Ask Clarity
  4. On-Device AI Reaches Production Viability

    Alibaba's Qwen 3.5 9B beats OpenAI's 120B model while running on 6GB RAM, Docker Model Runner enables zero-cost local inference with OpenAI API compatibility, and Apple put Apple Intelligence on a $599 device — collectively making on-device AI a production-ready option for cost-sensitive and privacy-sensitive features.

    Ask Clarity
  5. Team Dysfunction Diagnostics and the Coordination Bottleneck

    AI coding tools accelerate individual engineers while widening the shared-context gap, and Molly Graham's Waterline Model (structure → dynamics → interpersonal → individual) provides a diagnostic framework for the coordination dysfunction that results — making PM-driven decision documentation and context preservation the critical bottleneck, not engineering throughput.

    Ask Clarity

Deep Dives

Your Engineering Capacity Model Is Broken — Claude Code's 8-Month Takeover and OpenAI's GitHub Play Demand Immediate Recalibration

A rigorous 906-person survey of senior engineers (median 11-15 years experience) published this week delivers the most comprehensive picture yet of how AI coding tools have restructured engineering work — and the numbers should change how you plan your next quarter.

The New Operating Model

95% of engineers use AI weekly, and only 2.1% don't use it at all. But the headline number is more dramatic: 56% now do 70%+ of their work with AI, and 55% regularly use AI agents for code review, bug fixing, and automated tasks. This isn't autocomplete — it's a new operating model where engineers delegate entire workflows. Staff+ engineers are the heaviest agent users at 63.5%, and directors disproportionately favor Claude Code, meaning adoption is being driven top-down by your most senior technical leaders.

The Claude Code Disruption

Claude Code launched in May 2025 and became the #1 AI coding tool by February 2026 — dethroning GitHub Copilot, which had a 4-year head start. The mechanism is model quality: Anthropic's models are mentioned more than all other models combined for coding tasks. The tool is essentially a thin terminal wrapper around the best coding models available. Usage splits starkly by company size: 75% adoption in small companies vs. 56% GitHub Copilot dominance in 10K+ enterprises — a gap driven by 6-12 month procurement cycles that create a measurable productivity disadvantage for large organizations.

OpenAI's GitHub Competitor Changes the Platform Map

Simultaneously, OpenAI is building a GitHub alternative and actively discussing commercializing it. This isn't a research project — it's a product with a GTM motion. OpenAI's logic is clear: if AI-assisted coding is the future, owning both the AI models and the code platform captures the entire value chain. This directly fractures the OpenAI-Microsoft relationship and creates a new platform risk: if you've built CI/CD integrations, GitHub Actions workflows, or rely on GitHub OAuth, you're building on a platform about to face real competition.

The value in AI coding tools is migrating from the tool layer to the model layer. A single model release can reshape the entire competitive landscape — and 70% of engineers already use 2-4 tools simultaneously.

The Velocity Paradox

Here's the tension multiple sources surface: AI tools accelerate individual coding speed while doing nothing to preserve shared product context. Code changes that took 2 hours now take 2 days in mature systems — not because of technical debt, but because decision reasoning is buried in old tickets, Slack threads, and departed employees' heads. As one analysis frames it, throughput without alignment creates 'organized chaos.' Your PM role as keeper of product context becomes exponentially more important as AI-accelerated engineering produces AI-accelerated drift.


Cursor's $2B ARR (doubling from $1B in 3 months, 60% enterprise revenue, $29.3B valuation) validates that enterprises will pay aggressively for AI coding tools. This is now the benchmark your leadership will use to evaluate your AI feature adoption metrics. The sentiment gap between agent users (61% excited) and non-users (36% excited, 22% skeptical) is a management challenge — enablement, not mandates, is what shifts adoption.

What to do

  1. Audit your engineering team's AI tool usage against these benchmarks (95% weekly, 55% agent use, 63.5% Staff+) and identify adoption blockers by end of this sprint

  2. Fast-track Claude Code procurement/security review if your company hasn't approved it yet

  3. Map all GitHub integration touchpoints (APIs, Actions, OAuth, webhooks) and assess portability to alternative platforms by end of Q1

  4. Invest in decision documentation practices — ensure the last 10 major architectural decisions have discoverable reasoning, not just outcomes

The SaaS Pricing Crisis: Feature Moats Are Worthless, and the 99% AI Adoption Gap Is Your Biggest Opportunity

The Convergence That Should Worry You

Three independent analyses this week arrive at the same conclusion from different angles: the business model assumptions underlying most SaaS products are being repriced in real time.

From the cost side: AI has collapsed product build costs from ~$200K to as low as $20. Mid-tier AI models now match flagships at 40% less cost — Claude Sonnet 4.6 scores 79.6% vs. Opus 4.6's 80.8% on agentic coding benchmarks while costing $3/$15 vs. $5/$25 per million tokens. When a competitor can replicate your feature set in days for near-zero cost, your 6-month roadmap of functional improvements is a treadmill to nowhere.

From the demand side: OpenAI's own data reveals a devastating adoption gap. Only 50M of 900M weekly active users pay (5.5% conversion). Power users (95th percentile) use 7x more 'thinking capabilities' than median paid users. Roughly 2.5M people globally are extracting transformative value from AI. The other 897.5M use it as a slightly better search engine. OpenAI has decided to run ads on ChatGPT — signaling the subscription conversion model is underperforming.

The subscription model is narrowing to exactly two durable categories: utilities and continuously fresh context. If your product is neither, your pricing model is your biggest product risk this year.

Where Defensibility Actually Lives

Investors are now explicitly filtering for three criteria: unique data, deep workflow embedding, and autonomous task completion — not AI-assisted workflows. The shift from 'AI-assisted' to 'AI-autonomous' is critical: every feature that still requires a human to review, approve, or click through is a feature a competitor can leapfrog by removing that friction. Zillow's CEO is making this bet explicitly — the company's future isn't a better listings UI but transaction software that integrates across real estate agents and financing, growing during a housing crisis by going deeper, not wider.

The Adoption Gap Is Your Greenfield

The 99.75% of ChatGPT users who barely use it represent the largest warm market in tech history — 900M people who demonstrated intent to use AI but aren't getting value. The bottleneck isn't capability; it's workflow integration. The companies that win the next phase won't have the best models — they'll build the best bridges between AI capability and user workflows. Templates, guided interactions, progressive disclosure, domain-specific defaults. One analyst reports 30-50% time savings on editing, research, and translation when AI is properly integrated. If you can prove even a fraction of that for your users' specific jobs-to-be-done, you differentiate from competitors still selling on capability hype.


The 'Minimum Lovable Product' framework is gaining traction as the new launch bar. Most products sit at the functional and reliable layers. The ones winning retention and word-of-mouth invest in personality, delight moments, and human tone. Your next sprint planning should include: 'What's the emotional signature of this feature?' Meanwhile, Partiful displaced Facebook Events not through better features but by making every invite a growth loop — product-embedded distribution beating marketing spend.

What to do

  1. Model your revenue under three scenarios — current seat-based, usage-based, and outcome-based pricing — and present findings to leadership by end of Q1

  2. Measure your AI feature activation funnel: what % of users who encounter an AI feature use it in a way that delivers value (not just 'tried it once')? Compare your power-user vs. median usage ratio against OpenAI's 7x benchmark

  3. Build guided AI workflows (templates, suggested prompts, step-by-step wizards) for your top 3 user jobs-to-be-done this quarter

  4. Add a 'delight audit' to your sprint review — score each shipped feature on emotional resonance alongside functional completeness

Agent Infrastructure Is Forming a Real Stack — Your Product Needs an Agent-as-User Strategy, Not Just AI Features

The Stack Is Crystallizing Now

When you see prepaid virtual Visa cards for agents (agentcard.sh), email and phone capabilities (actors.dev), no-account hosting (iocaihost), session sharing (Droid), and on-demand app building (Tasklet's Instant Apps) all shipping in the same week, you're watching an infrastructure layer form in real time. This is the equivalent of watching AWS, Stripe, and Twilio emerge simultaneously — except compressed into months.

OpenAI's Stateful Runtime Changes the Architecture Decision

OpenAI and AWS are launching a 'stateful runtime environment' for enterprise AI agents within months. This is architecturally distinct from stateless API calls: agents maintain persistent memory about customers, business context, and conversation history. The AI infrastructure market is formally splitting into two tiers:

TierUse CaseCostExample
StatelessHigh-volume, simple API callsLowClassification, summarization
StatefulPersistent-memory agentsHighCustomer support, financial monitoring

OpenAI projects non-API/agent revenue will surpass API revenue by 2028. This is OpenAI moving up the stack from infrastructure to application platform — the classic play that historically crushes thin-wrapper products. The service circumvents Microsoft's exclusive stateless model distribution rights by selling agent services on AWS, not raw model access.

The GTM Stack Is Already Agent-Native

a16z's SR006 cohort reveals solo founders selling into banks, hospitals, and law firms with zero sales hires using browser agents, Clay/Lemlist pipelines, and AI-generated compliance docs. The emerging reference architecture:

  1. Browser agent (Claude Cowork, OpenAI Operator) → prospecting
  2. Enrichment (Clay, 11x) → data enrichment
  3. Sequencing (Lemlist, Instantly) → outreach
  4. CRM (Attio) → execution
  5. Compliance (Vanta) → trust artifacts

Notable absences: Salesforce, HubSpot, and Outreach don't appear once. One founder manages a literal 5-agent AI sales org over iMessage. CRM is being demoted to 'execution layer' — inference drives pipeline.

a16z predicts inbox saturation with AI-generated messages within 12 months. The outbound channel is degrading — trust artifacts, not outreach volume, are the real enterprise sales bottleneck.

Payment Rails Are Bifurcating

Agentic commerce is splitting into enterprise rails (OpenAI/Stripe, Google/Mastercard) for human-to-merchant flows and crypto-native rails (Coinbase x402, 50M+ USDC transactions, volume doubling monthly) for agent-to-agent commerce. For agents making thousands of micro-transactions per hour, card rails' 2-3% fees are a dealbreaker — stablecoins settle in seconds for fractions of a cent.

What to do

  1. Audit every critical user journey in your product and answer: 'Can an agent complete this programmatically?' Document gaps as your Q2 agent-readiness backlog

  2. Classify your current AI features as stateless-appropriate or stateful-appropriate; identify which would benefit from persistent memory and continuous operation

  3. Map the emerging AI GTM stack and identify where your product sits — layer, integration, or disintermediation risk

  4. Add 'trust artifact generation' (SOC 2 readiness docs, security architecture summaries) to your enterprise GTM feature backlog

The Waterline Model: Why Your Underperforming Team Is Probably a Structure Problem, Not a People Problem

A Diagnostic Framework From the Companies That Matter

Molly Graham — who spent two decades inside Google, Facebook, and the Chan Zuckerberg Initiative, and has since advised leaders at Stripe, Anthropic, OpenAI, Microsoft, and Gamma — presents the Waterline Model: a four-layer diagnostic for debugging team dysfunction. The layers must be worked in order:

  1. Structure — unclear goals, ambiguous roles, missing context
  2. Dynamics — leadership behavior, decision stability, process signals
  3. Interpersonal — relationship conflicts between leaders
  4. Individual — actual performance gaps

The mantra: 'snorkel before you scuba.' Graham's central claim is that a 'huge percentage' of team issues live at the structure level. Her case study: she took over a struggling marketing team, asked each person what their goals were, got wildly inconsistent answers, re-clarified the mandate and role definitions, and saw immediate performance improvement. No firings. No reorgs. Just structure.

Why This Matters More Now

This framework gains urgency in the context of AI-accelerated engineering. When 56% of engineers do 70%+ of their work with AI and individual throughput is skyrocketing, the coordination layer becomes the binding constraint. As one analysis frames it: code changes that took 2 hours now take 2 days in mature systems — not because of technical debt, but because decision reasoning is buried and undiscoverable.

Dynamics problems often show up as process issues but are rarely solved by process alone — they usually trace back to signals leaders send through their behavior.

The dynamics layer deserves special PM attention. Graham describes a founder constantly frustrated by slow team velocity — but the founder was swooping in to unmake decisions at the last minute. The result: the team adapted by adding extra alignment layers, escalating unnecessarily, and optimizing for not being wrong instead of shipping. If your team is slow and you can't figure out why, look at whether leadership behavior is creating an environment where speed is punished.

The Interpersonal Trap

Interpersonal conflict between leaders manifests as slowed decisions, hoarded information, and teams picking sides — but it often has structural roots (overlapping ownership, misaligned incentives). Only after ruling out structure and dynamics should you address relationships directly. At the individual level: evaluate the person against the role as it exists today, decide if the gap is coachable in the timeframe the business can afford, invest if yes, make a clean exit if no.

The meta-signal: even Anthropic, OpenAI, Stripe, and Microsoft are actively investing in organizational scaling support. If the most technically impressive organizations struggle with the human side, your team's coordination challenges are normal — but solvable with the right diagnostic.

What to do

  1. Run a Waterline audit on your most underperforming team: ask each member independently to state the team's top goal, their role, and how success is measured — if answers diverge, you've found your problem

  2. Add the four Waterline layers as a diagnostic checklist to your team retro or health check template

  3. Before your next performance conversation about a struggling team member, explicitly walk through structure and dynamics layers first

  4. Document whether leadership decision reversals are creating velocity drag — track specific examples over 2 sprints

The bottom line

The AI coding tool market flipped in 8 months (Claude Code is now #1, 56% of engineers do 70%+ of work with AI, Cursor hit $2B ARR), SaaS pricing models are breaking as build costs collapse to $20 and 99% of AI users can't extract value, and agent infrastructure is crystallizing into a real stack with payments, identity, and stateful runtimes shipping now. Your three highest-leverage moves this quarter: recalibrate engineering capacity against actual AI tool adoption, stress-test your pricing model against outcome-based alternatives, and audit every critical user journey for agent-readiness — because your next power users may not be human.