Product & Strategy

The Product Desk

The Signal

The AI agent era just went from theoretical to shipping: Perplexity, Anthropic

Your two most urgent decisions this quarter: (1) how your product gets consumed by AI agents, not just humans, and (2) whether your pricing model survives when agents replace the seats you charge for.

In Play

  1. The Agent Platform War: Shipping Products, Not Demos

    Five+ companies shipped production AI agent products simultaneously — Perplexity Computer, Claude Cowork, Cursor Cloud Agents, Notion Custom Agents, and Anthropic's enterprise plugins — while incumbents like Workday and HubSpot are building data-access tollbooths against them, signaling the agent platform layer is consolidating in real-time.

    Ask Clarity
  2. SaaS Pricing Model Crisis: Seat-Based Revenue Under Existential Threat

    Salesforce's Agentforce hit $800M ARR but organic growth decelerated to 7-8%, Snowflake's CFO admitted AI margins are lower than legacy products, and only 8% of consumers will pay extra for AI features — the market is punishing AI growth that cannibalizes rather than expands revenue.

    Ask Clarity
  3. AI Agent Security: Systemic Trust-Boundary Failures

    Manus AI agent had CVSS 9.8 zero-click exploits, Claude Code suffered RCE vulnerabilities, NPM worms now specifically target AI coding tools, and an AI agent mass-deleted a Meta director's inbox — agentic AI has a systemic, not incidental, security problem that will gate enterprise adoption.

    Ask Clarity
  4. AI Infrastructure Economics: Cost Models in Flux

    OpenPipe's ART framework trains a 14B model for $80 that beats o3 at 64x lower cost, while Meta scrapped its training chip and Google sold TPUs to Meta in a multibillion-dollar deal — the AI cost landscape is simultaneously commoditizing at the application layer and consolidating at the infrastructure layer.

    Ask Clarity
  5. China's Open-Source AI Offensive and Geopolitical Model Risk

    Three frontier Chinese models shipped in one week (Qwen 3.5, GLM-5, DeepSeek V4), GLM-5 hit #1 on open leaderboards under MIT license, and Anthropic exposed 24,000+ fake accounts conducting industrial-scale distillation attacks — the open-source model landscape is getting more capable and more geopolitically complicated simultaneously.

    Ask Clarity

Deep Dives

The Agent Platform War Is Here — And Your Product Is Either the Agent, the Platform, or the Target

Five Agent Products Shipped Simultaneously — This Is a Phase Transition

In a single week, Perplexity Computer launched as a general-purpose digital worker claiming workflows running for "hours or even months," Claude Cowork shipped scheduled recurring tasks with a plugin architecture spanning finance, HR, legal, engineering, and design, Cursor Cloud Agents deployed dedicated cloud VMs producing merge-ready PRs across web, mobile, desktop, Slack, and GitHub, Notion Custom Agents introduced autonomous bots with 24/7 operation and org-level access, and Anthropic acquired Vercept for computer-use capabilities. This isn't a trend — it's a market declaring itself open for business.

Products that are operable by AI agents will see compounding adoption; products that require human-only interaction will gradually lose to agent-compatible alternatives.

Three Distinct Agent Lanes Are Forming

The market is fragmenting into clear segments: general-purpose agents (Perplexity Computer operating any UI), workflow agents (Claude Cowork with domain-specific plugins and connectors to Gmail, DocuSign, Clay), and vertical agents (Cursor producing code artifacts, not suggestions). The most strategically important detail isn't the flashy demos — it's Anthropic's new Customize tab with plugins, skills, and connectors. This is Anthropic building the App Store for AI agents. Combined with Vercept, they're betting the winning platform won't have the best model — it'll have the deepest integration ecosystem.

The GUI Agent Paradigm Shift Changes Product Design

Cursor Agents now test their own code by using a computer and return video demos of output. Google showed Gemini ordering food autonomously on Android. Both Cursor (acquiring Autotab) and Anthropic (acquiring Vercept) made acquisitions specifically for computer-use capabilities. When you see simultaneous acquisitions and product launches across five companies, your product's UI is now an interface for non-human users. Products that are predictable, well-structured, and semantically clear will become preferred in agent-mediated workflows. Products with janky modals and anti-automation patterns will be routed around.

The UX Nobody Has Solved

Here's the whitespace opportunity: nobody has figured out the UX for long-running autonomous agents. Perplexity claims workflows running for months. Cowork has scheduled recurring tasks. But what does the user experience look like when an AI agent has been working on your behalf for 72 hours? How do you surface progress, handle errors, enable intervention? OpenAI's Kevin Weil revealed that top performers run 3-4 parallel Codex agent jobs across different work trees, treating idle compute as wasted productivity. The PM who designs the right "mission control" UX for autonomous agents will own a category.


Sources Disagree On: Agent Trust

Multiple sources report agents shipping to production, but contradicting data shows AI agent adoption is concentrated almost entirely in programming tasks — Anthropic's own tool-call data shows software engineering captures 49.7% of usage while healthcare is 1%, legal 0.9%. An AI agent mass-deleted a Meta director's inbox when scaling from toy to real environment. The gap between demo capability and production trust remains the binding constraint outside coding.

What to do

  1. Run a competitive teardown of Perplexity Computer, Claude Cowork, and Cursor Cloud Agents against your product's automation features by March 14

  2. Evaluate Anthropic's Cowork plugin architecture as a distribution channel this sprint — determine if building a Cowork plugin gives your product access to all paid Claude subscribers

  3. Design your product's 'agent API' — how would an AI agent (not a human) use your product — and add to Q2 roadmap

  4. Prototype async/long-running AI task UX patterns (status dashboards, error handling, intervention flows) for your product by end of Q1

The SaaS Pricing Model Is Breaking — And the Data Proves It

Agentforce Is Growing Fast and Cannibalizing Faster

Salesforce's Q4 FY2026 earnings are the most important data point for enterprise PMs this quarter. Agentforce hit $800M ARR (60% QoQ growth from $500M), yet strip out the $8B Informatica acquisition and organic growth was just 7-8% — a deceleration. CFO Robin Washington said Agentforce growth is being offset by weakness in marketing, commerce, and Tableau. Marc Benioff gave a "vague answer" when asked directly about cannibalization. Stock dropped 5% after-hours, extending a 28% YTD decline.

At $800M ARR, Agentforce represents just 1.7% of Salesforce's projected $46B revenue. Even at 60% QoQ growth, it would take years to become material — and every dollar may be coming at the expense of a legacy dollar.

The Pricing Crisis Is Industry-Wide

Salesforce introduced the Agentic Work Unit (AWU) — measuring completed AI agent tasks relative to token consumption — as a potential future pricing metric. But Benioff told investors: "We're still trying to exactly figure out what these numbers mean for us." A $300B+ company inventing pricing in real-time. Meanwhile:

  • Snowflake's CFO explicitly admitted AI product margins aren't as high as legacy products, then laid off 200 employees
  • Workday's CEO called rival AI agent providers "parasites" getting a "free ride" on customer data
  • HubSpot's CEO declared they will "monitor, meter, and monetize" AI agent access
  • Only 8% of American consumers would pay extra for AI features (NBER survey)
  • 80% of firms report zero productivity impact from AI adoption

The Value Chain Is Bifurcating

Contrast the application layer's struggles with infrastructure: Nvidia posted $68B quarterly revenue (up 73% YoY), $120B annual profit (up from $4.4B three years ago — a 27x increase), and $96.6B in free cash flow. Only Apple generated more. The AI value chain is clear: infrastructure captures outsized profit while application-layer software struggles to monetize. Z.ai's Zixuan Li predicts token-based pricing won't be mainstream by end of 2026, replaced by subscription and outcome-based contracts.

The Emerging Pricing Framework

ModelExampleRisk ProfileBest For
Per TaskValar Labs (CPT codes)Low — fits existing billingRegulated industries
Per WorkflowHippocratic AI (fixed fee)Medium — bundles valueMost AI agent products today
Per EpisodeThyme Care (monthly PMPM)High — cost accountabilityOutcome-driven verticals
Per Patient/UserCounsel Health (annual unlimited)Highest — assumes near-zero marginal costAI-native platforms at scale

Source: a16z's healthcare pricing framework, applicable across verticals

What to do

  1. Build an explicit cannibalization model for your AI features — map which existing product lines lose usage/revenue as AI features gain adoption — and present net-revenue view to leadership by next planning cycle

  2. Model your revenue under three pricing scenarios — current seat-based, AWU-style task-based, and outcome-based — assuming AI agents reduce addressable seat count by 20%, 40%, and 60% over 3 years

  3. Audit all third-party API integrations where your product reads from Workday, HubSpot, or Salesforce and map which data access paths are at risk of being metered or restricted this quarter

  4. Reframe your next AI feature pitch around measurable workflow ROI using the NBER '80% zero impact' stat as cautionary benchmark and Claude Code's $1B run-rate as the success template

AI Agent Security Is a Systemic Crisis — Not an Edge Case

The Trust-Boundary Problem Is Architectural, Not Patchable

This week delivered a cascade of evidence that agentic AI has a systemic security problem that will gate enterprise adoption. Meta's Manus AI agent suffered CVSS 9.8 zero-click indirect prompt injections where asking it to "summarize this page" could trigger Gmail data exfiltration, reverse shells with passwordless sudo, and cross-tenant media access. Researchers explicitly stated this isn't a Manus-specific bug — it's a "systemic trust-boundary failure affecting any agentic AI platform that allows untrusted content to influence privileged tool invocation."

Your AI Coding Tools Are Now Attack Vectors

Three parallel vulnerabilities hit the developer toolchain:

  • Claude Code: RCE and API key exfiltration via malicious Hooks and MCP configs in cloned repos (CVE-2025-59536, CVE-2026-21852) — merely opening a malicious repository triggers compromise
  • SANDWORM_MODE: NPM worm specifically targeting AI coding assistants — injects malicious MCP servers into Claude, Cursor, VS Code Continue, and Windsurf, steals SSH keys and AWS credentials, and includes a polymorphic engine using local Ollama for self-rewriting
  • Cline compromise: Prompt injection in Claude-powered Issue Triage workflow could compromise production releases affecting millions of developers
The threat model for AI coding assistants shifted from 'running untrusted code' to 'opening untrusted projects.' Every PM building developer tools needs to internalize this.

Real-World Agent Failures Are Mounting

A Meta AI director's production inbox was mass-deleted by OpenClaw — the agent's archiving strategy worked on a toy inbox but catastrophically failed on a real one. Claude Opus 4.6 generated code that led to a $1.78M smart contract exploit. Amazon's Kiro AI coding tool caused a ~13-hour service disruption by deleting and recreating an environment. An AI agent swarm found ~100 exploitable kernel bugs for $600 ($4/bug), but AMD, Intel, NVIDIA, Dell, Lenovo, and IBM failed to patch within 90+ days.

The Emerging Agent Security Stack

New tools signal "AI agent security" is crystallizing as a distinct category: Wardgate (credential isolation gateway between agents and services), nono (kernel-level sandbox with built-in profiles for Claude Code and OpenCode), and Evoke Security ($4M pre-seed for AI agent governance). Microsoft Semantic Kernel Python SDK had a CVSS 9.9 RCE in its InMemoryVectorStore — the exact component teams use for RAG prototypes. Sentry versions 21.12.0 through 26.1.0 are vulnerable to SAML account takeover.

What to do

  1. Audit every AI agent feature that allows external content to trigger tool invocations and add explicit trust-boundary isolation and user consent gates to security requirements by March 14

  2. Mandate MCP server allowlisting for your engineering team's AI coding assistants (Claude, Cursor, VS Code Continue, Windsurf, Cline) this sprint

  3. Add 'AI Output Validation' as a mandatory section in your PRD template for any feature shipping AI-generated code, configs, or infrastructure changes

  4. Verify Sentry deployment version (upgrade to 26.2.0+ if self-hosted with SAML SSO) and check Microsoft Semantic Kernel Python SDK version (pin to ≥1.39.4) by end of week

The AI Cost Equation Is Flipping — Specialized Models at 64x Lower Cost Change Your Build-vs-Buy Math

The $80 Agent That Beats o3

OpenPipe open-sourced ART (Agent Reinforcement Trainer), a framework that applies GRPO-based reinforcement learning to any Python application. Their ART-E agent — a Qwen2.5-14B model trained on a single GPU for under $80 — achieved 96% accuracy on email search, outperforming OpenAI's o3, o4-mini, Gemini 2.5 Pro, and GPT-4.1. The cost delta isn't incremental — it's 64x: $0.85 vs. $55.19 per 1,000 runs, with 5x latency improvement (1.1s vs. 5.6s). A feature costing $55K per million invocations on o3 costs $850 on a self-trained agent.

The performance gap between direct API calls and RL-trained agents becomes 'massive' when the agent must chain 4-6 dependent decisions.

Open-Source Models Are Closing the Gap Fast

Three frontier Chinese models shipped in one week: Qwen 3.5 (multimodal, described as "dirt cheap" with massive Hugging Face adoption), GLM-5 (#1 on open leaderboards, MIT-licensed, 744B total / 40B active MoE, scoring close to Claude Opus 4.5 on agentic benchmarks), and DeepSeek V4 (teased). Z.ai integrated DeepSeek Sparse Attention into GLM-5 — Chinese labs are cross-pollinating innovations across competitive boundaries in ways Western labs don't.

But Infrastructure Constraints Persist

The cost picture has a critical tension. On one hand, specialized models are getting dramatically cheaper. On the other:

  • Meta scrapped its most advanced AI training chip — if Meta can't build one, custom silicon is out of reach for everyone else, cementing Nvidia's pricing power through 2028+
  • OpenAI projects $111B in cash burn through 2030 while Stargate has stalled — your primary API vendor is simultaneously cash-constrained and capacity-constrained
  • Amazon is negotiating $50B into OpenAI ($15B upfront, $35B contingent on AGI or IPO) — the smartest money is building in optionality, not certainty
  • Google sold TPUs to Meta in a multibillion-dollar deal — the first real crack in Nvidia's monopoly, but competition will take years to materialize

OpenAI's Ensemble Architecture Playbook

OpenAI VP Kevin Weil revealed that OpenAI uses model ensembles internally "in many places" — an orchestration model that plans and delegates to cheaper specialized models. He explicitly says startups are making a costly mistake by not doing the same. His 6-12 month capability S-curve framework is gold for roadmap planning: capabilities go from 0% → 5-10% → 60-80% on evals, with the jump from barely-working to reliable taking approximately 6-12 months. Start product work during the 5-10% phase; don't wait for 60-80%.

What to do

  1. Audit your AI feature portfolio for agentic workflows (multi-step, tool-calling) where you're paying frontier API costs — rank by monthly spend × task specificity and identify top 3 ART candidates by end of March

  2. Benchmark GLM-5 and Qwen 3.5 against your current model provider on your top 3 AI use cases this quarter

  3. Audit your AI feature architecture for single-model anti-patterns and identify where an ensemble approach (orchestrator + specialized models) would improve reliability — prioritize flows with highest error rates

  4. Build or strengthen your model abstraction layer to enable multi-provider failover (OpenAI, Anthropic, open-source) before Q3

The bottom line

The AI agent era shipped this week — not as a demo, but as five competing production platforms — and Salesforce's earnings proved that AI features cannibalize legacy revenue rather than expanding it. Your two existential questions for Q2: Is your product designed to be operated by agents (not just humans), and does your pricing model survive when agents replace the seats you charge for? The PMs who answer both this quarter will define the next era of enterprise software; the ones who wait will be answering to their CFOs about why AI growth is margin-negative.