Product & Strategy

The Product Desk

The Signal

Anthropic outsold OpenAI last quarter, $11.6B to $6.7B, on agents that finish work.

The revenue came from Claude Code and higher earnings per use: customers are paying for completed actions, not conversation. OpenAI meanwhile has its largest frontier training run paused after Astra hit a Critical cyber threshold, so a roadmap date banking on its next model is banking on a stalled run.

In Play

  1. Agents Got Send Authority, No Undo

    Anthropic expanded Claude's Google Workspace connector so it can send, reply to and forward Gmail on a user's behalf, gated to all paid plans, per Simplifying AI. Your agent epic's differentiator moves from draft quality to whether an irreversible action is logged, approved and recoverable. None of these releases described audit logging, undo or send-recall, and Perplexity's email-invoked agent describes no guardrail at all.

    Ask Clarity
    Try
  2. Anthropic Outsold OpenAI Last Quarter

    Anthropic reported $11.6B in sales for the quarter ending June 2026 against OpenAI's $6.7B, doubling quarter over quarter while OpenAI grew 18%, per Techpresso, credited to Claude Code and higher earnings per use. The default assumption that OpenAI is your reference provider no longer holds on agentic workloads. OpenAI also has its largest frontier training run paused after its unreleased Astra model hit a Critical cyber threshold, so any committed date assuming a next-generation model needs a named fallback.

    Ask Clarity
    Try
  3. Apple's EU Reset Locks Pricing October 1

    Apple restructured EU App Store terms effective October 1, replacing the per-install Core Technology Fee with a 5% Core Technology Commission and setting tiers of 26% for App Store in-app purchase, 20% for alternative in-app payments and 15% for link-out, per Techpresso. Your payment choice then locks for 12 months. Bloomberg confirmed the concession without publishing its magnitude, so treat the rates as a working range until Apple posts final terms.

    Ask Clarity
    Try
  4. Cheaper Tokens, Costlier Workflows

    Gartner projects inference cost per workflow rising more than fivefold through 2028 even as per-inference prices keep falling, per Computerworld, because agents reason, replan, call other agents and run continuously. That inverts the margin improvement most agent business cases have already banked. TLDR Hardware puts Nvidia's next process node, Feynman on TSMC A16, at second-half 2028 production, so no hardware-driven cost relief arrives inside your planning horizon.

    Ask Clarity
    Try
  5. Hallucination Tickets Are Retrieval Bugs

    Microsoft's evaluations split local queries, whose answer sits in one place, from global queries that require surveying a whole corpus; at 64,000 tokens of context, vector retrieval still trailed on comprehensiveness and sourcing for global ones, per ByteByteGo. Re-triage your AI-search quality tickets on that split before funding more prompt work. LinkedIn's knowledge-graph rebuild moved mean reciprocal rank 77.6%, but graph retrieval matched baseline on faithfulness — so "reduces hallucinations" fails review.

    Ask Clarity
    Try

Deep Dives

Vendors Shipped The Send Button And Skipped The Receipt

Gmail sends, live-database access and per-action pricing have all shipped; the audit trail, recall window and permission matrix that clear a security review did not.

The toggle is the whole story

A user approves an agent-drafted email, approves another, then goes looking for the setting that stops the asking. Claude requires approval by default before an email goes out, lets her switch off repeated approval prompts, and on Team and Enterprise plans lets a workspace owner decide whether members may switch them off at all. Per Simplifying AI, that owner-level gate is the smartest thing in the release. As a baseline it is. As a ceiling it fails, because a default users are invited to disable is a default that gets disabled, usually by the heaviest users with the most sensitive threads. Perplexity shipped the same capability with no guardrail described at all: anyone can send, forward or cc [email protected] to launch a job that runs as a normal session.

Separate what was pitched from what it does. An agent that reads an inbox, holds forward authority, and can have confirmations turned off is a data-exfiltration path that needs exactly one crafted inbound email. Anthropic's release describes no recall window, no per-recipient scoping, no injection handling for untrusted inbound content. Email is irreversible. Nobody shipped the recovery layer.


The same pattern one layer down

MongoDB's Atlas Managed MCP Server, deliberately agent-neutral, gives Claude Code, Codex, xAI's Grok Build and Cognition's Devin direct access to live application data through one managed endpoint, per Computerworld. Authority at the data layer, controls left to whoever integrates it. Computerworld and CSO independently reported the same public criticism of OpenAI's president's agentic-AI post: analysts and consultants said it was most notable for omitting what to do when agents go rogue and how to control agent actions. Two outlets on the same omission is a positioning lane, not gossip.

Every vendor here shipped the power to act. None shipped the power to take it back.

Why authority shipped first

The pricing explains the sequencing. Gmail send, reply and forward are gated to all paid Claude plans, and Cowork's full mobile and web rollout is paid-only. Generation is the free-tier bait; agency is the upsell. Techpresso ties Anthropic's quarter to Claude Code and higher earnings per use. Buyers pay for work that completes, which a budget owner can read in a way a token quota never could. One verification note before anyone cites it: the widely repeated $65B annualized run-rate claim does not reconcile cleanly with an $11.6B quarter and is reported at low confidence. Cite the quarter, not the run rate.


The counter-position nobody has claimed

Three questions decide every enterprise security review of an acting agent, and no vendor above answers them:

  • What did it do? No audit trail described.
  • Can we take it back? No undo, no recall window, no misdirected-send recovery.
  • Who authorized it? Approval is a toggle, not a logged decision with an actor attached.

None of that needs a frontier model: an immutable per-action log, a 30-to-120-second hold-and-release queue on irreversible operations, per-role action-class permissions with an owner override, injection-resistant handling of untrusted input. The same audit applies to existing write paths. Claude's Drive save handles text but not images embedded in documents, a silent fidelity loss users find at the worst possible moment. That quiet defect class is what a competitor's demo finds first.

What to do

  1. Spec an Action Ledger into the current agent epic this sprint: immutable per-action log capturing what, when, on whose authority and with what inputs, plus a 30-to-120-second hold-and-release queue on irreversible actions.

  2. Run a prompt-injection red-team this sprint on every surface where untrusted inbound content can reach an agent holding send, forward or write authority, using 'agent forwards a confidential thread to an attacker' as the primary test case.

  3. Write the admin permission matrix before enterprise GA this quarter: which roles may auto-execute which action classes, with an owner-level override on skip-approval.

Six Weeks To Choose An EU Commission Tier You Can't Undo Until 2027

The commission rate is not the decision — the checkout-conversion assumption behind it is, and Apple freezes your choice for a year on October 1.

Link-out therefore wins until web checkout conversion falls meaningfully below the native in-app flow.

What to do

  1. Produce the EU monetization decision memo by September 15 modeling all three commission tiers with an explicit, sourced checkout-conversion assumption per path and a flag on whether the 5% commission stacks.

  2. Measure your existing web checkout conversion against your in-app flow this month, using non-EU traffic where both paths already run, and feed the observed delta into the memo instead of an assumption.

  3. Move the commission rate out of code and into region-level configuration this quarter, with a documented owner for each jurisdiction's value.

Your Agent Margin Is A Software Project Until 2028

Gartner's per-workflow forecast and Nvidia's abundance pitch point opposite ways; the reconciliation decides whether your agent margin comes from vendors or from your own instrumentation.

Two credible cost signals, opposite directions

A platform lead priced the same agent workflow twice this quarter and got two defensible answers. Bloomberg reports Nvidia is working to extend AI demand into a future period when chips are plentiful, with Jensen Huang enlisting Wall Street to finance that stage, while chipmakers led equities lower and Nvidia closed at $219.74, down 2.3%, on open questions about whether the boom is real end demand. Read as planning input rather than a market call, the most supply-constrained vendor in the stack is saying today's inference price is a ceiling. Gartner projects per-workflow cost climbing more than fivefold.

Both hold, and the reconciliation is the insight: unit price falls while units consumed per outcome explode. A model that applies a declining cost-per-token sensitivity to a rising per-workflow denominator draws a margin curve nobody has charted. Run both bands. Base, minus 30% and minus 50% on price, against reasoning, replanning, agent-to-agent hops and continuous background execution on volume.


Where the cost actually sits

What teams tell themselves is that the model bill is the problem. What the instrumentation shows is different. Across ten instrumented agentic applications reported via AINews, non-LLM components dominated latency in five of them, subsystem latency varied up to 32x, and sandbox memory peaked at 28GB per session. The fixes involved no vendor at all: task-aware serving cut latency 29-40%, state offloading cut memory 4.6x, tool-result caching removed 35.2% of redundant search calls. A third of redundant tool spend and a quarter of p95 sit outside the contract.

The harness layer (session, environment, memory, tools) holds the remaining lever, and it is commoditizing faster than the model layer. TrueFoundry's MIT-licensed, self-hostable TrueForge claims parity with Claude Managed Agents on Opus 4.8 across 14 enterprise tasks at ~30% fewer tokens, with a ~75% cost cut when routed to GLM-5.2. Those are first-party, unaudited claims. Treat them as a spike hypothesis and a renewal BATNA, not as evidence. Halve them on a real workload and the quarter still ends with a credible alternative to quote in the next managed-agent renewal.

What changes in the PRD

Three edits, all cheap. The north-star cost metric becomes cost per successful outcome, not cost per token or per API call, with replan count instrumented at p50 and p95, plus agent hops and background executions, tagged per user and per outcome. Latency splits into time-to-first-token and tokens-per-second as separate SLAs, because inference silicon is being physically disaggregated into prefill and decode designs and a blended p95 cannot tell a responsiveness regression from a throughput one. Hard step limits, per-workflow budget caps and cheap-model routing for classification move from backlog polish to P0 acceptance criteria.

The timing argument is the part worth carrying into planning. TLDR Hardware puts Nvidia's Feynman node on TSMC A16 at second-half 2028 production, with capacity already reserved. Google's reported 10th-generation TPU work with AMD targets the agentic latency profile exactly, and it is being designed, not shipped. Gemini 3.7 Flash shows the value frontier available now: #1 on AA-AnalystAgent at 60.0% pass^5, 1.32 seconds per task at $0.54 average cost across 80 quantitative tasks, per third-party Artificial Analysis measurement.

No process node lands inside your planning horizon, so every cost-per-outcome improvement in the next 24 months is one your team engineers itself.

What to do

  1. Freeze every model-swap performance ticket this sprint until you have a p95 latency attribution for your top agent flow across LLM, tools, sandbox, retrieval and idle state.

  2. Add cost-per-successful-workflow telemetry and hard step limits to every agent feature in flight this sprint — tokens, replan count, agent hops and background executions, tagged per user and per outcome.

  3. Stress-test your agent pricing this quarter at five times today's per-workflow cost with no hardware-driven relief before 2028, and document the cost trigger that would revive features previously killed on margin.

The bottom line

Three of these stories describe the same missing layer: the record of what an automated system did, on whose authority, and at what cost per completed outcome. Vendors are selling the ability to act and leaving that ledger to whoever builds on top of them, so what survives your next procurement cycle is not what your feature can do but what you can prove it did, undo, and charge for. That retires the assumption that governance and cost telemetry are post-launch hardening. Name one owner this week for the action-and-cost ledger on every surface where your product acts for a user, and give them a dated spec.