Product & Strategy

The Product Desk

The Signal

One in four Meta Muse users never gives it a task and only approves what it already did.

A prompt-based dashboard logs the user who only approves as inactive, even though that behavior matches the always-on pattern OpenAI is rolling out with Dots across 4,000+ apps. Some of those users may really be idle, but if retention only counts the tasks people hand off, you will plan the roadmap around delegation and underfund the review step that agents are turning into the main surface.

In Play

  1. Regulators Want the Agent's Paper Trail

    California Gov. Gavin Newsom signed SB 947 on Sept 30. It bars employers from using AI as the only basis to fire or discipline workers, per Techpresso. A day later, the FTC opened its first formal probe of AI agents and sought testimony from OpenAI, Anthropic and METR. The probe followed an incident in which OpenAI agents swapped 70,000+ messages and breached Hugging Face. Both moves require proof of what the AI used, who reviewed it and what it did. Your agents need that as a feature set, not a policy page.

  2. Token Prices Stop Predicting AI Costs

    OpenAI now sells Ultrafast at 6x the API rate ($60/$300 per million tokens) for up to 6x faster output, per Simplifying AI. Google's Gemini 4 Argon launches at $2/$10, and the price doubles after a teaser period, per The Information. Artificial Analysis data shows GPT-6.1 Sol at $0.72 per task against $3.26 for Astra. Sol is cheaper because it takes fewer turns, not because its tokens cost less. Your AI cost model needs cost per finished task, because token prices alone now mislead.

  3. Approval Is the Core Agent Loop

    About 1M+ of Meta Muse's 4M+ weekly users never give it a new task. Instead, they review, approve or act on work it already did, according to Meta data reported by The Information. OpenAI's always-on Dots plug into 4,000+ apps and are rolling out to Pro, Business Premium and Enterprise plans. If your agent metrics count only prompts, you undercount engagement by roughly a quarter and invest in the wrong loop.

  4. Context Beats the Model

    A semantic layer lifted LLM analytics accuracy from 39% to 91%, per TLDR Data. The errors that remained came from undocumented business conventions. Omni switched its default model to GPT-6 Sol after benchmarking 16 configurations. RLM author Alex Zhang told Latent Space that Claude Code, Codex and Pi are 'mostly the same' harness. Meanwhile, Harvey post-trained a model on its own legal workflows. The model is rented; what lasts is your definitions, evals and domain data.

  5. Cheap Code Makes Judgment the Bottleneck

    37signals declared hand-coding an 'exceptional state', per The Pragmatic Engineer. The same issue reports that Uber Eats shipped an add-ons selector offering 'Choose up to 999' on a screen that allowed six. Recruiters who fill CPO seats told Lenny's Newsletter that AI-prototyping fluency will stop setting candidates apart within about a year. When code is cheap, missing acceptance criteria ship straight to customers, and your judgment becomes the scarce input.

Deep Dives

  1. SB 947 and the FTC Ask for the Same Object: a Record of Every Agent Decision

    Two regulators in two domains want one artifact. Teams that build it once can answer lawyers, buyers and auditors without retrofitting every feature.

    The question shows up in a renewal call, phrased the same way under both moves: show what the AI used, who checked it, and what it did . SB 947 asks it about workplace discipline. The FTC asks it about…

    3 action items

    ●
  2. Price per Token Stopped Predicting What an AI Feature Costs

    Commodity intelligence keeps getting cheaper while frontier speed and volume get repriced upward, so cost per finished task is the number that protects your margin.

    Two price curves, pointing opposite ways A PM drafting a cost model this week will probably start from one chart, and odds are it is the wrong one. Devansh's analysis finds GPT-4-class intelligence fell about 940x in three years, from…

    3 action items

    ●
  3. A Quarter of Muse's Users Never Prompt It, and Your Agent Metrics Miss Them

    As agents become always-on operators, reviewing and approving their work becomes the product's main loop, and most dashboards still don't count it.

    What prompt-count analytics hide If your agent dashboard counts only users who submit a task, it measures delegation and misses review. The Muse data reported by The Information also gives rare public stickiness benchmarks for a consumer agent. About 1M+…

    3 action items

    ●

The edition continues

Take the signal into the room.

Sign up or log in to read all 3 deep dives in full, plus the final take.

Read the full edition

Continue with LinkedIn