Product & Strategy

The Product Desk

The Signal

Four states took Meta to trial asking a court to delete infinite scroll and autoplay.

Design-defect theory has won three straight, and Section 230 has stopped working as a dismissal argument. What's on the table is $1.4T plus a nationwide remedy, which is the part that reaches teams nowhere near this docket: every retention mechanic a minor can reach, Stories and beauty filters included, is now discoverable evidence of a design decision rather than a growth lever. The engagement experiment you scope this quarter comes with a paper trail attached.

In Play

  1. Engagement Mechanics Go On Trial

    Four state attorneys general — California, Colorado, Kentucky and New Jersey — opened trial against Meta in Oakland on Aug 18. Morning Brew reports the remedy list names mechanics: infinite scroll enabled by autoplaying video, ephemeral content like Instagram Stories, and beauty filters. That turns any retention loop minors can reach from a growth lever into a discoverable design decision. The Information notes Section 230 has repeatedly failed to get these design-defect claims dismissed.

    Ask Clarity
    Try
  2. Cost-Per-Task Becomes The Sales Weapon

    Glean is publicly benchmarking $0.45 per task against $1.84 for Claude Cowork, per Latent Space's interview with CEO Arvind Jain, and credits the gap to harness and routing rather than model quality. Per-user enterprise AI spend is up 10–20x year over year: agentic runs got much longer while frontier token rates only rose 2–4x. Computerworld reports boards now ask what already-spent AI money returned before funding more, so cost-per-task belongs in your success metrics, not your infrastructure notes.

    Ask Clarity
    Try
  3. A Free Laptop Model Ties A Frontier Tier

    Alibaba's Qwen3.8-27B — free to download, multimodal across text, image and video, and able to run on a laptop — scored level with OpenAI's GPT-5.6 Luna on the Artificial Analysis Intelligence Index and passed 1 million downloads in days, The Information reports. Luna is the tier most teams route high-volume, low-complexity calls to. So both your cheap cloud line item and your 'we can't do on-prem' answer to stalled security reviews are now product decisions. Chinese-origin weights make provenance review the remaining blocker.

    Ask Clarity
    Try
  4. Agent Absorption Is A Churn Vector

    SaaStr canceled Notion after seven years without any competitor winning a bake-off: an internal AI agent gradually absorbed the last workflow the subscription served. Usage and health-score models structurally cannot see that pattern, so the renewal risk arrives as gentle softening rather than a cliff. Vercel separately holds first-party data on agents creating accounts and buying third-party products through its Marketplace, per its CEO's session with a16z speedrun — a buyer most funnel analytics still count as human.

    Ask Clarity
    Try
  5. Vendors Claim The Layer Above The Model

    Anthropic shipped /design into Claude Code and Desktop as a research preview: it generates editable artboards, then implements the selected one as working code. Cursor launched Origin, Git hosting with bidirectional GitHub sync. ElevenLabs moved agent creation, config edits and pre-change LLM cost estimates into an MCP server that runs inside Claude. Nvidia entered model routing with NeMo Switchyard. Routing layers and vendor dashboards are now weak places to spend roadmap capacity.

    Ask Clarity
    Try

Deep Dives

The Remedy List Reads Like A Hostile PRD Review

Three consecutive plaintiff-favorable outcomes taught state attorneys general to ask for feature deletions instead of fines, and this ask would apply nationwide.

The precedent chain is what makes this one different

The safety lead who pulled her app's under-18 numbers this week did it because of a docket, not a roadmap review. Design-defect claims against consumer software are now three for three. A Los Angeles jury found Meta and Google liable in March for one young woman's depression and anxiety, awarding $6M on a defective-product-design theory. A New Mexico judge then ordered $942M in total penalties, $567M on top of an earlier $375M, called Meta a public nuisance comparable to air pollution, and mandated in-state safety features. Thousands more suits are pending. Meta's CFO Susan Li told investors in July that this year's trials "may ultimately result in a material loss." CFOs do not use that phrase about nuisance litigation.

The damages figure is the part that gets briefed upward and the part that matters least. The presiding judge called the states' $1.4 trillion demand "unreasonable" at a pretrial hearing, then called Meta's own $4M estimate "also unreasonable" in the same breath. Reuters puts the states' internal figure closer to $200B. What actually ships out of this trial is injunctive relief against named mechanics, and unlike New Mexico's in-state order, this one would potentially apply nationwide. That collapses the state-by-state compliance arbitrage most consumer teams have quietly relied on.


Two of the remedies are builds, not deletions

Separate what the remedy list is called from what it requires. Plaintiffs want Meta forced to detect minors running multiple accounts and to build parental verification for underage users. Those are staffed engineering programs with schedules, not config flips, and if granted they become de facto national requirements. Meta's courtroom posture makes the exposure map worse for everyone else: its lawyer argued Facebook is overwhelmingly used by adults while conceding Instagram skews younger, and the narrative pushes Snapchat forward as the app better known for appealing to kids. That is risk transfer. Whoever skews youngest without defensible age assurance inherits the target, and no one can make the counterargument without knowing under-18 and unknown-age share by surface.

MechanicPlaintiff framingWhat you need on file
Infinite scroll + autoplayRemoves natural stopping cuesSession-break controls; documented age gating
Ephemeral content and timersExpiry manufactures compulsory returnNo expiry-triggered pings to minors
Beauty and appearance filtersHarms adolescent body imageDefault-off for under-18, labeled, adult-gated
Streaks and loss-aversion loopsEngineered compulsionWritten user-benefit rationale; minor default state
Time-spent as an OKRIntent evidence in discoveryReplace with task-completion or retained-value

MIT Technology Review's coverage of an unrelated case supplies the checklist that generalizes. Flock Safety's roughly 120,000 license plate readers, where the harm traced back to four choices: what to collect, who can search it, how long it is retained, how widely it is shared. Misuse controls arrived retroactively. A fifth row belongs on that list, which is abuse cases reviewed before GA.

Section 230 covers what your users publish. It does not cover how you engineered the loop that keeps them publishing.

Five independent reports read the same trial and reach the identical conclusion, which is unusual: the money is theater and the injunction is the product event. The live disagreement is scope. Appeals could narrow a nationwide order, and Meta's demographic defense may genuinely undercut intent. The sort worth running in the next planning cycle has two axes. One is cost to build proactively. The other is cost under subpoena. Everything sitting in the cheap-now, expensive-later cell goes into the sprint, and the scope uncertainty does not move a single item out of it.

What to do

  1. Inventory every engagement mechanic on minor-reachable surfaces now — auto-advancing feeds, autoplay, streaks, expiry timers, appearance filters, notification cadence — and attach a dated, written user-benefit rationale to each.

  2. Instrument and report under-18 and unknown-age share of DAU by surface within this sprint, and stop citing assumptions in risk conversations.

  3. Scope age assurance, parental consent and multi-account detection as one feature-flagged platform epic this quarter, with two vendor quotes plus a build estimate.

Cost-Per-Task Is Now The Slide That Wins The Deal

Vendors stopped selling capability and started selling unit economics, while a free laptop-sized model removes the last technical excuse for a single-provider cost structure.

The mechanism behind the cheaper number

An admin opened the model-selection panel, read the list, and left it on automatic. Glean says customers do that overwhelmingly, "mostly because of cost." That choice is the tell for what is actually being sold. The claimed advantage is routing sequence, not model choice. Waldo, the agentic search model launched in April, decides how to break down a question, which tools to use, what to read next, and when it has enough evidence to hand off to a frontier model, so model selection happens after the context is assembled. Arvind Jain's framing: "We're able to assemble the raw materials needed to do the work without burning LLM tokens." Claimed payoff: 50% lower latency and 25% fewer tokens. Two adjacent patterns are becoming table stakes. One is three-tier model control, where end users pick a model, admins restrict models and set usage limits, and an automatic mode routes per task. The other is that refusing to call an LLM at all now ships as a feature; Jain's example is users asking the product to multiply two numbers.

Separate the thing being pitched from the thing being measured. The $0.45-versus-$1.84 comparison, the 4x claim against Claude Code, and Waldo's 50% and 25% figures are vendor-reported without disclosed task mix or methodology. They are good evidence that the competitive axis moved. They are not numbers that survive a slide review. The valuation carries its own caution: $300M ARR at $7.2B is roughly 24x, which holds only if the budget compression Jain describes doesn't bite.


The sources disagree on the direction of token prices

The Information reads Alibaba as commoditizing the low end: a free 27B model that ties OpenAI's cheapest flagship variant and satisfies air-gapped deployment. Latent Space reports the same inflection from the buyer side, with open-weight usage moving from "minuscule" to essential inside roughly three months, at about an order of magnitude cheaper per task. Both arrows point down.

Paul Smalera's reporting points up, and it is the less comfortable read. Nvidia licensed Groq's technology and hired founder Jonathan Ross with much of his team, then invested in what was left at $3.5B, roughly half its $6.9B September 2025 mark. Groq now resells inference capacity instead of designing chips. The credible price competitor in inference silicon was neutralized without an acquisition. The planning assumption is flat-to-rising token costs on the frontier tier while the open-weight floor keeps dropping, because that is what the two stories describe together.

MetricWhat most teams reportWhat buyers now ask for
QualityBenchmark or accuracy scoreCost per correct answer
CostAggregate monthly inference spendCost per completed task, p50 and p90, by intent
LatencyMedian response timep95, including reasoning overhead
ReturnPrompt volume, DAU touching AIPayback period the customer can verify in their own billing

TheSequence sharpens the latency half: a feature needing sixteen samples plus a majority vote pays roughly 16x per query forever, and its p95 disqualifies in-editor, voice and real-time surfaces entirely. Compounding Quality supplies the investor frame boards already use, which treats personalization as a moat only while unit cost stays flat as engagement rises. Engagement is a weak stand-in for value, but it is the axis the chart gets drawn on anyway, so chart cost-per-personalized-session by engagement decile before someone external does. The forcing function for this sprint is narrow. Every shipped feature names its cost-per-task and its p95. Anything that cannot produce both numbers is a prototype.

Enterprise AI stopped competing on what the model can do and started competing on what the task costs. A team that cannot state its cost-per-task cannot defend its roadmap.

What to do

  1. Instrument cost-per-completed-task at p50 and p90, segmented by intent class, and put it beside task success rate in every AI PRD this sprint.

  2. Run a five-day bake-off routing your two highest-volume, lowest-complexity AI calls to a self-hosted Qwen3.8-27B, then take the resulting on-prem SKU to your three most stalled security reviews.

  3. Add a documented gross-margin floor per AI feature this quarter, stress-tested at plus-40% inference cost with a named mitigation for each: model downgrade path, caching, rate limit, or price change.

Notion Lost Two Customers, For Two Different Reasons

One departure names three fixable defects in any retrieval feature; the other is invisible to every health score you run, and together they redefine who your product's real user is.

Churn one: a power user, three named defects

Casey Newton had thousands of saved links in Notion and gave up on its agentic search. He replaced it with a prompt written by Claude Fable 5, pasted into a terminal, writing Markdown into Obsidian, seeded with his entire archive since 2020. One month later the stack holds 1,440+ pages, including a single Meta page running 12,000+ words with 1,300+ internal links, auto-generating concept pages and refreshing a home page every morning. His job title contains no engineering.

He never filed a model complaint. The three reasons he left read like a PRD checklist: in-database search was keyword-only, the agent lived on a separate app surface from the data, and it failed to cite sources without extra prompting. The second failure in the same account travels further. His manual concept-tracking system died of proliferation. Nothing was removed; the marginal cost of maintenance outran the marginal value. Products that ask the user to be the librarian have already priced that decay into the retention curve. He calls the replacement the most useful thing he tried this year and too clunky to recommend, which is the textbook precondition for a commercial winner. Score the managed-LLM-wiki category (Town, CEO Jean-Denis Greze) before funding an internal build.


Churn two: nobody's dashboard moved

The SaaStr cancellation came with no competitive loss event, and the analysis says plainly that traditional usage and customer-health metrics stop predicting anything under agent absorption. The detectable fingerprint is a ratio, not a level: programmatic, API and MCP read volume rising per account while human seat sessions fall. Retro-test it against the last four quarters of churn. If it would have flagged half of them a quarter earlier, it has paid for itself.

The same reporting supplies the vocabulary buyers bring to renewals. Google shipped least-privilege agent identities, flow-level access controls, audit trails and human approvals into the Workspace admin console. Docker began logging AI governance decisions and streaming them to SIEM. Teleport shipped short-lived database certificates for agents touching production data. Dynatrace paid ~$915M for Arize to join model output and agent behavior to app traces. Four vendors converging on one control list means the list is a procurement checklist now, not architecture advice.

The buyer may not be a person

Vercel shipped an unpromoted CLI capability and saw an immediate usage uptick. Agents reach new products faster than its own employees absorb them. It holds data on agents creating accounts and buying third-party products through a Marketplace that Guillermo Rauch describes as "almost like an Amazon.com for the cloud." Alipay launched a full agentic commerce foundation plus the AHA cross-device protocol suite, turning authorization, fulfillment and settlement into a protocol across thousands of already-adapted services.

Read the magnitude carefully. Vercel's usage claims are unquantified, "an uptick," "data showing," and a16z discloses that third-party information was not independently verified. Direction is high-confidence, size is unproven. What the sources agree on is narrower and worse. The failure is instrumentation, not capability. None of these accounts left because a model underperformed.

An agent evaluated, signed up for, and bought software. Analytics that cannot tell it apart from a human are optimizing the funnel for the wrong user.

What to do

  1. Ship an agent-adjacency retention indicator this sprint: machine read volume per account plotted against human seat activity, retro-tested on four quarters of closed churn.

  2. Audit every AI retrieval surface you own this sprint against three requirements: hybrid semantic-plus-keyword search inside the primary data surface, citations default-on with no extra prompting, and no separate agent tab.

  3. Tag signups, CLI and API calls, and purchase events by client signature this quarter, and report agent share of activations as one number in roadmap review.

The bottom line

These stories rhyme on one shift: judgments a product manager used to make privately are now graded by outsiders — a courtroom reading design rationale, a finance owner reading unit economics, and a customer's own automation deciding whether your interface is needed at all. Intent no longer stays internal, and margin is no longer an engineering detail. Both are discoverable, and both get quoted back at you by someone holding leverage. Rewrite your PRD template so every retention mechanic carries a written user-benefit rationale and every AI surface carries a unit-cost ceiling with a named owner.