Product & Strategy

The Product Desk

The Signal

Anthropic doubled Claude Code enterprise pricing the same week it launched a $1.5B PE

This splits your market in two: PE-backed companies will get Claude mandated top-down before your sales call arrives, while your Claude-dependent features face a pricing squeeze that makes the 17x-cheaper DeepClaude alternative a necessity, not an experiment.

In Play

  1. PE Firms Become AI's Distribution Channel

    OpenAI's 19-firm Wall Street consortium and Anthropic's $1.5B JV with Blackstone/Goldman/H&F deploy AI into thousands of portfolio companies via top-down mandate. Blackstone alone carries 250+ companies. Your enterprise sales motion is being bypassed — one deal with a PE sponsor replaces hundreds of individual sales cycles.

    Ask Clarity
  2. AI Feature Costs Explode for Power Users — Breaking Flat-Rate Models

    GitHub Copilot's $40/month subscription consumed $221 of inference in 15 agentic messages. Uber burned its entire 2026 AI coding budget in 4 months at $500–$2K/engineer/month. Agentic workflows trigger 10–50x compute vs. chatbot queries. The top 5% of users generate half the inference spend — flat pricing is arithmetic, not strategy.

    Ask Clarity
  3. Harness Engineering: The Moat Is in Context, Not the Model

    Changing only the prompt/middleware harness moved GPT-5.2-codex from 52.8% to 66.5% on Terminal-Bench 2.0 — a 13.7-point lift without touching model weights. OpenAI has formalized 'harness engineering' as a discipline. Context overload (burying 400 tokens in 200K) is the #1 agent failure mode. Teams investing in model selection are investing in the weekend-swappable part; the 6-month part is context pipeline.

    Ask Clarity
  4. AI Prototyping Reaches Production Quality — Design Systems Are the Key

    Stripe built Protodash — AI prototyping connected to its Sail design system via MCP — and PMs now use it equally with designers. Generic AI tools produce 'blurple slop' but design-system-aware tools produce artifacts stakeholders argue about as if they were real products. The meeting changed from 'should we staff a designer' to 'here's a working demo, how do we improve it.'

    Ask Clarity
  5. AI Capability Doubling Every 7 Months — Roadmap Horizons Compressing

    Autonomous task duration: 30 seconds (2022) → 12 hours (2026) — a 1,440x increase in 4 years. SWE-Bench: 2% → 93.9% in 2.5 years. ClawMark benchmark simultaneously shows multi-day agent tasks still fail across all frontier models. Features scoped 12+ months out face a moving capability floor; features scoped for this quarter hit a reliability ceiling.

    Ask Clarity

Deep Dives

Your Pipeline Is Being Pre-Empted: PE Firms Are Now AI's Distribution Layer

The New Buying Motion Isn't Bottom-Up Anymore

A PM at a mid-market SaaS company noticed three of her top twenty target accounts changed ownership fields this week. All three sit under Blackstone portfolio companies. The deals stalled not because a competitor appeared but because the buyer committee was replaced overnight by someone with a portfolio-wide vendor consolidation mandate and a spreadsheet that does not care about her champion's product love.

Seven independent intelligence sources point at the same pattern this week. OpenAI closed a $10B raise from a 19-firm Wall Street consortium explicitly designed to push ChatGPT agents into every mid-market company those firms own. Five days later Anthropic closed its $1.5B JV with Blackstone, Goldman Sachs, and Hellman & Friedman under the same structure. Blackstone alone carries 250+ portfolio companies. The full consortium runs into the thousands.

For three years the labs sold direct through enterprise sales and API contracts. Now they sell once to a PE sponsor and deploy to hundreds. This is Accenture-style distribution at venture speed.

What This Actually Means For Your Product

The motion has two faces that compound against independent vendors:

  1. Top-down mandate: When Blackstone owns a company and says 'implement Claude for operations efficiency,' the company implements Claude. The outside sales call arrives after the decision is made.
  2. Consultant-bundled deployment: Anthropic is not selling an API anymore. It is selling a transformation program with Claude inside. OpenAI, Anthropic, and Salesforce have each built consulting arms for the same reason: enterprise buyers do not want a model, they want someone to run the workflow change.

The commercial implications diverge sharply by segment. PE-owned accounts move top-down the moment the sponsor writes the mandate into the operating plan. They do not run bottom-up experiments, and a third-party tool is not going to win on developer love. Independent companies still pick on time-to-value and usage depth. One motion does not serve both.

Where You Can Still Win

The PE JV deploys general-purpose AI across operations. It does not deploy domain-specific workflow tools that require proprietary data and context. The 2x2 that matters: on one axis, is the product a layer the foundation model provider will eventually ship as part of a deployment package, or a layer they will keep routing customers toward because it makes their consulting engagements faster? Build in the second cell. The products that survive are consultant-resellable and priced per-outcome. They slot into the JV's deployment playbook rather than competing with it.


The Timeline Is Quarters, Not Years

Goldman Sachs is already cutting Claude access for Hong Kong bankers over contract concerns while simultaneously being part of the $1.5B JV. The internal contradictions tell you the playbook is still forming. The window to position as complementary rather than competitive is the next 6-12 months, before the sponsor's standard vendor checklist and operating-plan template harden around a specific deployment pattern across hundreds of companies.

What to do

  1. Map your customer base against PE consortium ownership (Blackstone, Goldman, Hellman & Friedman, General Atlantic) — identify which accounts are now inside the OpenAI/Anthropic distribution lock-in

  2. Redesign your enterprise pricing to include a portfolio-level conversation option — priced for a buyer who compares line items across 8 companies at once

  3. Brief your VP Sales on the PE-as-distribution shift and propose a partner motion: position your product as the domain layer that makes Claude/GPT deployments more valuable inside specific verticals

The $221 Problem: Flat-Rate AI Pricing Is Breaking Under Real Usage

The Numbers That Should Change Your Next Pricing Meeting

A developer named Theo sat down with GitHub Copilot and ran 15 agentic messages. One of those messages consumed more than 60 million tokens. By the end he had burned $221 of inference tokens on a $40/month subscription. Uber, running the same playbook at scale, disclosed Claude Code costs of $500–$2,000 per engineer per month and burned its entire 2026 AI coding budget in four months. This is not misuse. This is the power-user behavior the product was designed to enable.

The distribution of cost per user in an AI product is not normal. It is a long tail where the top 5% of users generate roughly half the inference spend. A flat subscription is a bet that the average covers the tail. On current model costs, it does not.

Why Agentic Workflows Break Every Cost Model

The compute multiplier is the structural problem. A chatbot query generates one model call. An agent workflow generates 10–50x that in tool calls, retries, context assembly, and evaluation loops. The Atlantic documented a full reversal in data center sentiment from 'too much capacity' to infrastructure panic, pinned specifically on agents rather than foundation models.

Workload TypeCost Per InteractionMultiplier vs. Chat
Simple completion$0.01–$0.051x
RAG query$0.10–$0.505–10x
Agentic coding session$5–$15100–300x
Multi-day autonomous task$50–$200+1,000x+

Anthropic, reading the same tea leaves, doubled Claude Code enterprise token costs and is optimizing for high-margin enterprise over developer adoption. Meanwhile DeepClaude now offers "identical autonomous loops" at 17x lower cost by swapping Claude's backend for DeepSeek V4 Pro. The spread between what users pay and what inference costs is widening in both directions.

The Pattern That Works

Microsoft's shift to consumption-based pricing is not thought leadership. It is a pricing architecture change from the company generating more SaaS revenue than any other on Earth. The shape: seats as packaging for prepaid consumption, with overage billing per token, per agent action, or per outcome. Replit's 300% net revenue retention and positive gross margins are the existence proof. They own the full stack and sell to non-technical users who do not price-compare against API bills.

Here is the 2x2 for this sprint. One axis: is the product priced on model cost, or on user outcome? The other axis: is the model a dependency, or a substitutable component? Products in the 'outcome pricing + substitutable model' cell survive both the subsidy unwind and the Jupiter/DeepSeek resets. Products in the 'cost-plus + single-model dependency' cell get repriced by someone else's strategy.

What to do

  1. Stress-test every AI cost line item at 3x current projections — specifically model agentic sessions where one user interaction triggers 10+ model calls

  2. Run a pricing architecture workshop this sprint to model hybrid seat + consumption pricing — identify every AI interaction that could be metered (tokens, agent actions, outcomes)

  3. Ship cost-per-user telemetry before the next pricing meeting — segment users by inference cost and identify whether your top 5% are customers you want or customers you're subsidizing

The Harness Is the Product: Where to Build Your AI Moat This Quarter

A 13.7-Point Improvement Without Touching Model Weights

Mason Drxy demonstrated that changing only the prompts and middleware in a coding agent harness moved GPT-5.2-codex from 52.8% to 66.5% on Terminal-Bench 2.0. Anthony Maio stated the implication directly: lock-in comes from 'how repo state is fetched, ranked, and compressed into the prompt,' not from the shell or the model. A second researcher claims >20x cost reduction from tuning open models inside well-designed harnesses.

OpenAI has now formalized 'harness engineering' as a distinct discipline, with four named capabilities agents need from their environment: queryable runtime states, progressively-disclosed documentation (AGENTS.md), merge-first CI, and background cleanup agents. This is the first time a frontier lab has described the wrapper as more important than the model.

A thin wrapper around an API with basic retrieval has no moat and is one LangGraph upgrade from commoditization. Domain-specific context compression — knowing which pieces of a user's data matter for a query and how to rank them into a prompt — is defensible IP.

Context Overload Is the #1 Production Failure Mode

Teams shipping agents this quarter consistently report the same finding: the most common reason agents fail is too much context, not too little intelligence. A typical agent has a 200K token context window. The actual instructions it needs are ~400 tokens, buried under tool definitions, reference docs, and brand guides. The SKILL.md pattern — two clear routing lines per skill file — is what's working in production, letting the LLM route without embeddings or retrieval layers.

The Investment Framework

Separate three layers and invest accordingly:

  1. Model layer (commodity): Swapping models is a weekend of work. No moat here — leadership rotates monthly across Anthropic, OpenAI, Google, and DeepSeek.
  2. Context pipeline (defensible IP): How you fetch, rank, and compress user data into prompts. This is 6 months to rebuild and where engineering investment belongs.
  3. Orchestration layer (watch): Sakana's 7B Fugu model hit SOTA on GPQA-Diamond by learning communication topologies via RL — signaling that hand-engineered orchestration DAGs may get displaced by learned policies within 18 months.

Multi-Model Is Now Mandatory

The Claude-vs-GPT character split is a segmentation signal: users pick Claude for drafting and GPT for structured output. Products assuming one model fits both lose half their segment. With MCP reaching 10,000+ public servers and OpenAI, Google, and Anthropic all speaking the same protocol, switching costs between providers trend toward zero. The harness that routes dynamically per task — not the model choice — is where retention lives.

What to do

  1. Audit your context pipeline this sprint — document exactly how you fetch, rank, and compress context into prompts and assess whether this is competitive advantage or liability

  2. Assign a named PM to own the AI harness — with a roadmap, eval suite, and budget for context-engineering work

  3. Architect your AI abstraction layer so swapping the underlying model requires a config change, not a migration — target completion before Q3

The bottom line

The AI product market split into three layers this week and your pricing, distribution, and engineering strategy need different answers for each: PE firms now control AI distribution into thousands of mid-market companies (map your pipeline against their portfolios before Q3), flat-rate AI pricing is mathematically broken when power users consume $221 of inference on a $40 subscription (ship cost-per-user telemetry before your next pricing meeting), and the durable moat isn't the model — it's the context pipeline that a 13.7-point benchmark lift from harness changes alone just proved is worth more than model selection.