Product & Strategy

The Product Desk

The Signal

Nikhyl Singhal's data from hundreds of PM career conversations confirms the split is

In the same week, Claude's tokenizer silently inflated your inference costs 12-27% without a pricing email, Clay/Figma/PostHog committed to two-track billing (seats for humans, consumption for agents), and Flask creator Armin Ronacher documented 'vibe slop' shipping to production across 30+ engineering teams.

In Play

  1. The PM Role Is Splitting — Only Builders Survive

    Coordination PMs are being eliminated structurally. Builder-PM roles are at multi-year highs with comp up. The hottest archetype isn't 'product manager' — it's the domain builder (sales builder, HR builder, finance builder) who ships without a team. PM-to-engineer ratios heading to 1:12.

    Ask Clarity
  2. The 60% Bundling Threat + Stealth Cost Shifts

    Platform vendors now ship 60% AI versions of point solutions — good enough to kill $80K contracts. Annual deals hide the churn. Meanwhile, Claude's tokenizer silently raised effective costs 12-27%, agent tasks burn 100-1000x chat-level tokens, and two-track billing (Clay, Figma, PostHog) is the new pricing standard.

    Ask Clarity
  3. Vibe Slop Goes Enterprise — AI Code Quality Crisis

    Flask creator Armin Ronacher documented code quality decline across 30+ teams as AI-generated 'vibe slop' ships to prod. Kent Beck names it the 'Genie Tarpit' — AI code scores low on both correctness and flexibility, creating compounding debt. PMs are part of the problem: AI-generated rebuttals are overriding senior engineer gatekeeping.

    Ask Clarity
  4. Your AI Moat Moved: From Models to Verifiers and Domain Skills

    A 4.2M-parameter verifier improved LLaDA-8B from 22% to 60.7% on math reasoning — without changing the base model. Amazon's COSMO rejects 65-91% of raw LLM output before it ships. a16z found structured domain 'skills' boosted agent success from 10% to 70%. The moat is the verification layer, not the model.

    Ask Clarity

Deep Dives

The Builder PM Mandate: Coordination Work Is Being Permanently Eliminated

The Career Ladder Just Inverted

Nikhyl Singhal, drawing on hundreds of career conversations through his Nikhyl.AI platform, published the most uncomfortable analysis of the PM profession this year. Open PM roles are at a multi-year high. Compensation is up. Both facts are true — but only for hands-on builders. A 20-year product veteran at Amazon-caliber companies, laid off in 2023, is still searching 2+ years later. The market isn't slow. It's hiring a different PM.

The old career ladder — builder IC → management → coordination of builders — spent twenty years looking like career growth. AI tools have absorbed the substrate of coordination work: meeting notes, status rollups, stakeholder translation, cross-functional alignment documents. Singhal calls this shift structural, not cyclical, meaning the coordinator-PM role is not coming back when the market recovers.


The New Archetype: Executive Builder

The winning profile has three components: function depth, product mindset, AI fluency. Missing any one moves a candidate into a smaller pile. The hottest emerging archetype isn't labeled 'product manager' at all — it's the 'sales builder,' 'HR builder,' 'finance builder' who can systematically obsolete manual work inside traditional business functions.

Executives are actively hunting for product-minded people who can ship against deep domain knowledge without a team. That reframes every hiring decision in the room this quarter.

Corroborating data from TLDR Founders recommends 4-5 person squads with the PM directly evaluating output quality, not relying on metrics through layers of project management. PM-to-engineer ratios are heading to 1:12 at companies that have taken the last eighteen months seriously. Chainguard now expects engineering managers to be at the 50th percentile of token usage among their direct reports — treating AI tool proficiency as a management competency, not optional upskilling.


The Diagnostic That Matters This Week

Singhal's forcing function is a simple 2×2. One axis: does the work produce an artifact a customer touches, or coordinate the people who produce it. Other axis: is the work something a model plus one builder can do by Friday, or does it require sustained human judgment over weeks. The only safe cell is 'produces artifact + requires sustained judgment.' Every other cell is under pressure on a twelve-month horizon, and the 'coordinates people + model can do it' cell is already being cleared.

AI is removing the shipping bottleneck that disadvantaged non-technical PMs, making deep domain knowledge the scarce asset rather than technical ability. Years spent understanding enterprise sales, healthcare operations, or financial compliance got more valuable — but only when paired with the ability to actually build.

What to do

  1. Audit your last two weeks: count artifacts that touched a user vs. artifacts that touched another employee. If the ratio is worse than 1:3, restructure this sprint.

  2. Build functional proficiency with AI coding tools (Cursor, Claude artifacts, Replit) to the point where you can prototype features solo by end of Q3.

  3. If you manage PMs: redesign team structure to eliminate pure coordination roles by next planning cycle. Replace with builder-ICs who own end-to-end delivery.

Your Product's 60% Problem — Platform Bundling and Stealth Cost Shifts Are Squeezing Simultaneously

The Renewal Conversation Changed

A CRO sits on a renewal call with one question: why are we paying $80K/year for this when the bundled copilot does most of what we need? Platform vendors are shipping AI-augmented 60% versions of specialized feature sets. 60% is not parity. 60% is enough to make the buyer demand a discount instead of a product — which is a different business with worse margins. HighLevel is the proof point: one platform replacing 10+ tools, $5.2B in facilitated sales, 7M+ AI voice calls.

Annual deal structures are hiding the churn. The renewal dashboard reads green. The decision to leave was made months ago.

The evaluation criteria shifted underneath the renewals. Buyer questions moved from 'Does it integrate with Salesforce?' to three new gates: Can agents drive this product? Are the APIs clean? Is there an MCP connector? Fail any one and there's no rejection email — the buyer simply moves on.


Stealth Cost Shifts Make It Worse

Anthropic's Claude Opus 4.7 tokenizer change is what cost drift looks like when nobody sends a pricing email. The sticker price didn't move. The new tokenizer improves understanding but inflates effective costs 12-27% for typical inputs. Teams running RAG pipelines, summarization, or document analysis on Claude just absorbed a quiet COGS increase. This is a this-sprint issue — finance will notice one to two billing cycles late, which is the worst possible time.

Simultaneously, agentic workloads burn 100-1000x more tokens than chat. A single Claude Code bugfix consumes ~900K tokens. METR data shows autonomous task horizons doubling every 131 days — from 4 minutes on GPT-4 to ~12 hours on Claude Opus 4.6. Features priced on chat assumptions are money-losing products.


Two-Track Billing Is Now the Standard

Clay, Figma, and PostHog have committed to two-track billing systems — seats for humans, consumption for agents. These aren't AI-native startups experimenting. They're established, product-led companies restructuring billing infrastructure because AI agents are a meaningful share of their 'users.' Any product with an API surface inherits this question. The 80/20 routing pattern emerging from Mendral is the architecture response: cheap Haiku handles 80% of routine tasks, expensive Opus handles the rest — and costs went down after upgrading to a frontier model. This tiered approach fundamentally changes which AI features are economically viable.

What to do

  1. Recalculate COGS for every Claude-dependent feature using the 12-27% tokenizer inflation. Flag any feature that goes margin-negative and present to finance this sprint.

  2. Run the '60% audit': for your top 10 accounts, identify which features a platform competitor could replicate at 60% with AI augmentation. Calculate revenue at risk and present to leadership.

  3. Model agent-native pricing scenarios: what happens when 20%, 40%, and 60% of your API consumption comes from AI agents? Design two-track billing architecture before agents arrive at scale.

  4. Prototype Mendral's 80/20 routing pattern for your highest-cost AI feature. Route routine tasks to cheap models and reserve frontier models for complex orchestration.

Vibe Slop Goes Enterprise: 30+ Teams Report AI Code Quality Decline — And PMs Are Making It Worse

The Problem Has a Name Now

A senior architect on one of the thirty-plus teams Armin Ronacher surveyed rejected a pull request last month as unnecessary complexity. The junior who authored it pasted the rejection into an AI agent and shipped back a polished rebuttal in ten seconds. That is the 'vibe slop' Ronacher, creator of Flask and now at Sentry, is seeing across more than thirty engineering teams. These are not experimental setups. They are production environments where AI-generated code is moving past review processes calibrated for human-speed output.

Kent Beck named the diagnostic in a parallel analysis. His 'Genie Tarpit' thesis says AI code generators produce output that scores low on both correctness ('does it work?') and flexibility ('can we change it later?'). The flexibility cost is invisible for weeks or months. Then teams hit what Beck describes as it's not a gentle slowdown. It's a wall. In his words: 'Complexity piles on complexity until even the genie can't pretend to make progress any more.'


PMs Are Part of the Problem

This is the part PMs should sit with longest: juniors and PMs are using AI-generated counterarguments to override senior engineer pushback. The architect who rejects an approach now spends thirty minutes dismantling a rebuttal that took ten seconds to generate. The cost of producing a bad argument has collapsed. The cost of refuting one has not. Run that loop across every technical decision in a sprint and the most experienced engineers get exhausted into compliance. Teams tell themselves this is 'alignment.' What is actually happening is that seniority is being rate-limited by token throughput.

The honest self-check for any PM is whether they have used AI to make a shaky product case sound more technically defensible. If so, that is the dynamic.

DX Is the Prerequisite, Not the Feature

CircleCI data says 90th-percentile DX teams ship 2x+ faster with AI tools than they did before. Teams with legacy codebases, slow CI, and scattered documentation are not capturing the same lift. AI makes engineers on teams with good developer experience faster, and leaves everyone else roughly where they were. The gap is compounding.

Beck's admission that 'nobody knows' how to solve the tarpit is the signal for dev-tool PMs. He lists six speculative approaches and endorses none. That is a named, authoritative, unsolved problem with engineering leaders as the buyer. For PMs not building dev tools, the decision on Monday is a process one: audit AI-generated code for architectural coherence, not just functional correctness. Put two numbers on the same dashboard. features merged per week AND time-to-diagnose the next production incident in that surface area. If both rise together, the roadmap is borrowing from a future sprint.

What to do

  1. Audit your team's AI-assisted code review: measure review depth (time spent, comments per PR) on AI-generated vs. human-written code over the last 60 days. Present findings at next retro.

  2. Establish explicit team norms: when senior engineers reject an approach, AI-generated rebuttals must be flagged as such. Add this to your team's working agreement.

  3. Stress-test Q3/Q4 roadmap commitments against a scenario where AI-assisted velocity degrades 30-50% due to accumulated complexity. Present the risk-adjusted timeline.

  4. Track features merged per week alongside time-to-diagnose-next-incident as paired metrics starting this sprint. Report both to leadership together.

The bottom line

The PM profession split into two jobs this week and only one of them is hiring: Singhal's data shows builder-PM demand at multi-year highs while a 20-year Amazon veteran searches for 2+ years, Ronacher documented AI 'vibe slop' degrading code quality across 30+ production teams (with PMs making it worse by weaponizing AI-generated rebuttals against senior engineers), and your unit economics silently shifted as Claude's tokenizer inflated costs 12-27% without a pricing email while Clay, Figma, and PostHog committed to the two-track billing architecture (seats for humans, consumption for agents) that makes your per-seat model obsolete. The PMs who survive 2026 are the ones who can ship artifacts users touch, catch the quality decline AI is creating, and model their feature economics under pricing structures the industry already adopted.