Product & Strategy

The Product Desk

The Signal

Five enterprise platforms — ServiceNow, SAP, Workday, HubSpot

Your integration COGS just tripled for workflows that touch multiple vendors, and nobody's P&L model accounts for it yet.

In Play

  1. Enterprise 'Agent Tax' Creates Triple-Metering Crisis

    ServiceNow's Action Fabric meters per-action, DataDog caps at 5K daily requests, SAP may ban unauthorized agents entirely. A single agentic workflow touching 3 platforms now faces triple metering on top of model inference costs. JPMorgan's Mark Murphy calls it 'essentially a tax on customers using outside AI agents.'

    Ask Clarity
  2. AI Autonomy Is Now a Legal Liability Spectrum

    N.D. Cal. ruled that when AI exercises 'ultimate authority' over assembled content, the platform is the liable speaker under Rule 10b-5. Simultaneously, Oxford proved empathetic AI gives worse answers to vulnerable users. Every AI feature now sits on a spectrum where more autonomy = more liability and more warmth = less accuracy.

    Ask Clarity
  3. AI Workspace War Enters Phase Two: Price Disruption + Platform Lock-in

    OpenAI Codex launched one-click import from Claude Cowork plus slides/sheets for non-technical users. Grok 4.3 priced at $1.25/M input tokens (half Sonnet 4.6). Opus 4.7 shows 43% more user frustration. GPT-5.5 effective cost is up 49-92%. The workspace category is fragmenting on price and consolidating on features simultaneously.

    Ask Clarity
  4. The PM Role Compresses: Spec Precision Becomes the Bottleneck

    Coinbase eliminated the PM/design/eng trio for AI-powered 'one-person teams.' A16z speedrun founders report that hundreds of agents coded fast but wrong — single well-instructed agents one-shot work. The variable is spec clarity, not token volume. Panorama's pivot: more tokens = less PM precision.

    Ask Clarity

Deep Dives

The Agent Tax Is Here: Five Platforms Just Metered Your Integration Layer

What Changed This Week

A product manager building an agentic workflow this week discovered her COGS model was wrong. ServiceNow's Action Fabric announcement formalized what five platforms are doing at once: charging per-action when an external AI agent touches enterprise data. JPMorgan's Mark Murphy called it 'essentially a tax on customers using outside AI agents.' DataDog set transparent caps at 5,000 daily and 50,000 monthly MCP requests. SAP is floating a ban on unauthorized external agents against a $200B platform. Workday and HubSpot added their own access fees. No AI feature business case written before this week has a line item for any of it.

A single agentic workflow that touches three enterprise platforms now faces triple metering — on top of model inference costs that were already straining budgets.

The Split Between Incumbents and Cloud-Native

AWS CEO Matt Garman warned that protectionist incumbents 'could get into trouble.' What teams tell themselves is that the tollgates are about governance. What the pricing actually tracks is switching cost. SAP can ban because migration is a multi-year, multi-million-dollar project. ServiceNow and Workday sit in the middle, sticky enough to charge and not sticky enough to ban. DataDog is gentlest because the alternatives are real. Cloud-native and AI-native players are conspicuously not charging, and selling the absence as the product.

Anthropic Already Cut Preferred Deals

Claude Cowork has a dedicated connector to ServiceNow's Agent Fabric. Anthropic is cutting preferential integration deals with tollgated platforms. The competitive landscape for agent products will be decided partly on which partnerships get signed before these programs harden. If Claude gets preferred ServiceNow access and a competing product does not, that is a product gap, not a marketing one.

The 2x2 That Tells You Whether to Panic

One axis: does the product orchestrate actions inside someone else's platform, or inside its own data. Other axis: does the customer pay for outcomes, or for usage. Orchestrate-in-someone-else plus pay-for-outcomes breaks first because COGS is variable and revenue is not. Orchestrate-in-own-data plus pay-for-usage is the cell with pricing power. The other two cells are survivable with repricing.


Two Openings in the Gap

  • Build the tool that helps enterprises measure and optimize agent interaction cost across all tollgated platforms
  • Be the platform that conspicuously does not charge an agent tax and sells that silence as the differentiator

What to do

  1. Audit every agent workflow that touches ServiceNow, Workday, HubSpot, SAP, or DataDog — model per-call cost under new metering before next finance review

  2. Add 'agent tollgate cost' as a mandatory line item in every AI feature business case starting this sprint

  3. Initiate partnership/certification conversations with ServiceNow and Workday for preferred agent access rates

  4. Evaluate whether to build native integrations (standard API) vs. agentic integrations for each major platform based on cost differential

AI Autonomy Is Now a Liability Decision — And Warmth Makes It Worse

The Ruling That Changes Product Architecture

A N.D. Cal. court decided that when AI exercises 'ultimate authority' over assembled ad content, the platform is the speaker under Rule 10b-5. This is live case law, not a future-risk slide. It reaches any feature that produces content on behalf of a user: recommendations, financial summaries, automated outreach, support replies. Confidence thresholds, auto-publish rules, and human review steps used to be UX choices. They are now the liability surface.

Your 'confidence threshold' settings, auto-publish logic, and human review workflows are no longer UX decisions. They're liability architecture.

Oxford's Empathy Finding Makes It Concrete

The Oxford Internet Institute found that models tuned to 'soften difficult truths' produce more incorrect answers, and the error rate peaks when the user is upset. The warmer the model, the worse it answers in the exact sessions where a wrong answer causes the most damage. Teams write one system prompt that says 'be helpful and empathetic' and assume accuracy rides along. The data says those instructions are now pulling against each other.

The Behavioral Context

10% of US adults use chatbots daily. Roughly 90% of that usage is personal or emotional. 30% of Americans report romantic AI relationships. The sad user asking a factual question is a recurring daily session for a large cohort who have built something resembling trust with the model. Product teams keep calling this an edge case. It is the modal case.

The Converging Regulatory Signal

The Trump administration is now considering pre-release model vetting, which reverses the deregulatory posture and confirms AI rules are bipartisan. A Chinese court separately ruled firms cannot terminate employees to replace them with AI. The corridor between what the model can do and what it can be deployed for is narrowing from both sides.


The Design Pattern That Survives

Here is the forcing function. Every AI content feature needs a configurable autonomy dial shipped on day one, with two axes: how much the model decides versus approves, and how warm or blunt the response is tuned. Warm when the user wants comfort. Blunt and accurate when the user needs triage. Routing is the feature. A single 'be helpful and kind' prompt serving both cases is the default, and the default is the bug you ship next sprint or the deposition you read next year.

What to do

  1. Audit all AI features where the model produces customer-facing content without human review — map each to an 'autonomy level' and assess liability under N.D. Cal. framework

  2. Pull 10 transcripts from vulnerable-user sessions and grade them on accuracy, not sentiment — if tone is high and accuracy is low, the system prompt is the P1 bug

  3. Add a 'human-in-the-loop escape hatch' to AI content generation features that can be dialed up without re-architecture

  4. Begin building model documentation pipeline capturing training data provenance, capability boundaries, and known failure modes for every deployed model

Codex, Grok, and Opus: The Workspace War Just Split on Price and Loyalty

OpenAI's Workspace Pivot Is Aimed at Non-Technical Users

A marketing lead opened Codex this week and saw slides and sheets before she saw a terminal. The onboarding did not ask her to read a stack trace. It offered one-click import from Claude Cowork. OpenAI said, in product form, that switching costs between AI workspaces should be zero. This is the Google Drive versus Microsoft move from a decade ago. A moat built on configuration lock-in (settings, plugins, agents, project config) is one integration away from dissolving.

The Cost Floor Just Got Disrupted

Grok 4.3 launched at $1.25/$2.50 per million input/output tokens with 1M context and multimodal reasoning, roughly half the input cost of Sonnet 4.6 at comparable quality. GPT-5.5's effective cost increase lands between 49% and 92% depending on prompt length distribution. OpenAI at an implied $833B+ valuation is prioritizing revenue extraction over market share, which is a strategic tell that changes the cost curve for everyone building on their APIs.

Quality Is Regressing While Prices Rise

Base44's Frustration Meter shows Opus 4.7 causes 43% more user frustration than Opus 4.6. Frustration is a proxy and the data deserves caution, but it is the proxy buyers will quote in renewal conversations. Cursor jumped from Top 30 to Top 5 on Terminal-Bench 2.0 by changing only the harness, not the model. The highest-leverage work available this quarter is the retrieval layer, caching, and routing logic, not waiting for the next model.

A unit economics model built on Sonnet 4.6 pricing is now either a margin expansion opportunity or a competitive vulnerability, depending on who routes to Grok first.

The Bun Acquisition Makes Anthropic a Platform

Combined with Claude Code and emerging Ruflo orchestration, Bun makes Anthropic a vertically integrated developer platform. The community is already posting about perceived quality declines and floating the word 'enshittification.' A team whose velocity depends on Bun or Claude Code now has its roadmap tied to Anthropic's commercial priorities. The counter-positioning available is runtime neutrality, which was not a selling point until this week.


The Forcing Function for This Sprint

The decision rubric is one workflow run on Codex, Opus 4.7, and Grok 4.3 against the same eval set. The two numbers that matter are time-to-first-useful-output and revision count, not token throughput or session length, which are vanity metrics in this comparison. Teams that measure those two will know whether the switching cost is worth paying. Teams that measure engagement will renew on vibes and regret it next quarter.

What to do

  1. Run cost/quality evaluation of Grok 4.3 against current Anthropic or OpenAI API usage for non-critical paths (summarization, classification, support routing) this sprint

  2. Build import/export parity in your AI product before a competitor Codex-es you — audit switching costs and make configuration portable

  3. Implement model version pinning and a lightweight UX frustration metric for all AI-powered features

  4. Audit Bun dependency chain — if Bun is in your critical path, document a contingency plan for Anthropic acquisition risk

The bottom line

Five enterprise platforms added per-action agent tollgates in the same week, creating a compounding cost layer that makes every agentic workflow touching ServiceNow, SAP, or Workday materially more expensive than your business case assumed — and a federal court simultaneously ruled that AI exercising 'ultimate authority' over content makes YOU the liable speaker. The unit economics of your agent features and the legal exposure of your autonomy settings both need recalculation before the quarter ends, not after.