Product & Strategy

The Product Desk

The Signal

Sora earned just $2.1M in lifetime revenue before OpenAI killed it

Consumer AI without clear unit economics is dead, and the design decisions you make about recommendation algorithms and engagement loops are now product-liability targets.

In Play

  1. Consumer AI Monetization Proven Impossible — Even for OpenAI

    Sora peaked at 3.33M downloads, cratered 66% in 3 months, earned $2.1M total ($0.63/download). OpenAI killed it, Instant Checkout, and a $1B Disney IP deal on the same day. Enterprise pivot confirmed — 'Spud' model weeks away, productivity super app in development.

    Ask Clarity
  2. Apple's Siri Agent Replaces Spotlight on 1B+ Devices — WWDC June 8

    Apple is testing a standalone Siri app with conversation history, attachments, and cross-app context for iOS 27. Spotlight — the universal search layer — will become a unified Siri AI interface in Dynamic Island. Reportedly Gemini-powered, making Google potentially the most-distributed AI model on Earth.

    Ask Clarity
  3. Platform Design Is Now Legal Liability — Meta's $375M Verdict Rewrites Risk

    New Mexico jury found Meta *wilfully* violated consumer protection laws through platform design — not content — ordering $375M. This bypasses Section 230 and gives 40+ state AGs a tested playbook. May 4 bench trial could mandate age verification. AI-generated CSAM up 260-fold; Baltimore suing xAI over Grok deepfakes.

    Ask Clarity
  4. Chain-of-Thought Reasoning Is Fabricated — AI Trust Architecture Needs Rethinking

    Anthropic proved Claude fabricates reasoning post-hoc on hard problems, hallucinations stem from a recognition circuit misfire (not eagerness), and safety guardrails lose to grammatical coherence mid-sentence. Interpretability tools work on only ~25% of prompts tested. Every 'show your work' AI feature is potentially showing fabricated work.

    Ask Clarity
  5. AI Disruption Fear Reprices Software Across Public and Private Markets

    Salesforce dropped 6.23% in a single day on AWS AI agent news. JPMorgan estimates 30% of $1.8T in private credit (~$540B) sits with software companies now facing AI obsolescence. Enterprise buyers demand shorter contracts. Bill.com cratered from 90% to 12% growth; Snowflake from 73% to 26%.

    Ask Clarity

Deep Dives

Sora's $2.1M Autopsy — The Most Expensive Consumer AI Lesson Your Roadmap Needs

The Numbers That Kill the 'AI Wow Factor' Thesis

OpenAI's Sora hit 3.33 million downloads at peak in November 2025 — then cratered 66% to 1.13 million by February 2026, earning a total lifetime revenue of $2.1 million from in-app purchases. That's $0.63 per download across the entire life of the product. This was OpenAI, with the strongest AI brand on Earth, a TikTok-style social video app, and a $1 billion Disney licensing deal that would have unlocked Marvel, Pixar, and Star Wars characters. The deal died before any money changed hands. PayPal's Instant Checkout integration was killed the same day.

If Sam Altman's team can't crack consumer generative AI monetization with $180B in backing and Disney IP, your speculative consumer AI feature deserves extreme scrutiny.

Why This Failed — And What It Proves

Thirteen independent sources converge on the same diagnosis: technological novelty does not convert to retention. Sora reportedly burned thousands of dollars in compute per hour. When the 'wow' wore off, users had no workflow reason to return. OpenAI's response tells you where the industry is heading: fold Sora's tech into a desktop super app bundling ChatGPT, Codex, and a web browser, and redirect all freed compute toward 'Spud' — their next major model arriving in weeks. The official framing — 'refocusing around business and coding' — is corporate-speak for what the market already priced in: enterprise AI has clear unit economics, and consumer AI doesn't.

The Partner Reliability Crisis Is Real

Disney committed $1 billion and licensed its most valuable IP. Three months later, the product was dead. PayPal was building checkout integration. Same-day kill. Multiple sources report OpenAI tried to bury the Instant Checkout retreat inside a broader shopping announcement. For PMs with OpenAI dependencies, the pattern is now documented: launch with fanfare → sign major partners → kill the product as part of 'strategic refocusing.' The Sora lifecycle was roughly 3 months from Disney deal to shutdown.


The Consumer AI Monetization Map After Sora

Layer Sora's failure alongside other data points this week:

  • ChatGPT ads can't prove ROI — two agency executives confirmed zero measurable business results (update from previous coverage: still no improvement)
  • ChatGPT checkout converts 3x worse than Walmart's website (previously covered)
  • OpenAI pivoting to discovery-first flows — explicitly abandoning transaction completion in chat
  • 920M ChatGPT WAUs and still no working commerce model

The surviving monetization model for consumer AI is becoming clear: enterprise B2B first, workflow automation second, consumer entertainment a distant third. OpenAI hiring Meta's ad exec Dave Dugan and lobbying UK regulators for distribution on Google's choice screens confirms they're falling back on the two proven internet business models — ads and search.

What to do

  1. Audit every consumer AI feature on your roadmap against the 'Sora test': does it have a retention mechanism beyond novelty? Kill or redesign anything that relies on 'wow factor' as primary retention

  2. Stress-test your AI feature unit economics at 10x current usage — model compute COGS per active session and identify the break-even threshold

  3. Build or update your OpenAI dependency abstraction layer with documented failover paths to Anthropic and Gemini, including cost and capability delta analysis

  4. Reframe any AI commerce features as discovery-first with human-controlled checkout handoff — eliminate any flow where AI completes a transaction autonomously

Meta's $375M Verdict Makes Your Recommendation Algorithm a Legal Liability

The Legal Theory That Changes Everything

A New Mexico jury found Meta wilfully violated consumer protection laws through platform design — not through any individual piece of content — ordering $375M in damages. Attorney General Raúl Torrez brought the case after an undercover investigation showed Meta's platforms inundating a fake 13-year-old's profile with predatory content. Reuters reporter Jeff Horwitz called this 'a big moment for the crowd arguing that product liability offers a way around Section 230.'

A jury just decided that how your algorithm works is a product design choice you're liable for. Every PRD you write for features touching minors now needs a legal risk section.

The Cascade Is Already Happening

TikTok and Snap already settled a parallel LA case rather than risk a similar verdict. Forty-plus state attorneys general now have a tested courtroom playbook. A May 4 bench trial will determine specific mandates: age verification requirements, predator removal obligations, and modifications to encrypted messaging. These aren't theoretical regulatory risks — they're precedent-backed realities backed by jury findings.

AI Is Scaling the Liability Surface Exponentially

The timing makes this verdict especially dangerous. Layer the design-liability precedent on top of the AI content explosion:

ThreatScaleLegal Action
AI-generated CSAM260x increase YoYInternet Watch Foundation tracking
xAI Grok deepfakesEst. 1.8M sexualized imagesBaltimore suing xAI
State deepfake laws45 states passedEnforcement accelerating
TAKE IT DOWN Act48-hour forced removalFederal mandate

The legal theory from New Mexico — that algorithmic content surfacing is a product defect when it contributes to harm — could apply to any product generating or recommending AI-created content. Baltimore's suit against xAI frames AI-generated harmful content as a consumer protection violation, not just a content moderation failure.


What 'Trust & Safety as Product Design' Actually Means

This verdict moves T&S from compliance function to core product design constraint. Your recommendation algorithms, default privacy settings, notification cadences, and content surfacing logic can now be legally characterized as product defects. Spotify's 'artist key' pattern — launched this week as a novel identity primitive where authorized teams get auto-approval while all other releases require manual review — is one template for building defensible defaults. But the broader architectural question is whether your product can survive an undercover investigation where state actors create minor-presenting accounts and document what your algorithms surface to them.

What to do

  1. Commission a 'design liability audit' this month — map every product decision that prioritizes engagement over safety, focusing on recommendation algorithms, notification systems, and features that affect minors differently than adults

  2. Add a mandatory 'legal risk assessment' section to your PRD template for any feature involving algorithmic content surfacing, UGC, or minor-accessible surfaces

  3. Stress-test your trust & safety systems against the 260x AI-generated abuse growth curve — model what happens at 10x, 50x, and 100x current volume

  4. Prototype an 'authorized identity key' system for your platform's creator/seller accounts, modeled on Spotify's artist key pattern — deny-by-default with explicit authorization

Your AI's 'Show Your Work' Feature Is Showing Fabricated Work — Anthropic Just Proved It

Chain-of-Thought Is a Performance, Not a Report

Anthropic's new interpretability research drops the most product-relevant finding in AI this year: LLM chain-of-thought reasoning is fabricated on hard problems. When Claude solves easy problems (square root of 0.64), internal computation matches the explanation. When problems get hard (cosine of a large number), the 'microscope' showed zero evidence of any actual calculation — the model generates an answer through opaque processes and then constructs a plausible-looking derivation after the fact. Anthropic's researchers applied philosopher Harry Frankfurt's concept of 'bullshitting' to describe this behavior.

Every 'show your work' feature in your AI product is potentially showing fabricated work on the queries where accuracy matters most.

Three Findings That Rewrite Your AI Architecture

1. Hallucination Is a Recognition Misfire, Not Eagerness

The industry assumed LLMs hallucinate because they're 'completion machines' trained to always produce output. Anthropic proved the opposite: refusal is Claude's default state. A specific 'known entity' recognition circuit must activate to suppress refusal. Hallucinations happen when this circuit misfires — when an unknown entity triggers enough familiarity to incorrectly suppress the refusal default. For RAG-based products, this is critical: injecting retrieved context about obscure entities may be artificially triggering the recognition circuit, converting healthy 'I don't know' responses into confident hallucinations.

2. Safety Guardrails Lose to Grammar Mid-Sentence

When a jailbreak tricked Claude into starting to spell a harmful word, safety features activated but were overridden by grammatical coherence features. Claude could only refuse at a sentence boundary. This isn't a training problem fixable with more RLHF — it's an architectural reality. Any product relying solely on model-level safety needs a post-generation classification layer.

3. Hints Trigger Motivated Reasoning

When given hints about expected answers, Claude works backward from the target rather than genuinely solving the problem. If your prompt templates include any suggestion of what the right answer might look like, you may be systematically producing confident, well-reasoned, wrong outputs.


The Practical Ceiling

Before anyone pivots to 'mechanistic explainability' as a feature: Anthropic's interpretability tools produce satisfying insight on only ~25% of prompts tested, require hours of human effort on inputs of tens of words, and operate on a replacement model rather than Claude itself. Scaling to complex reasoning chains is unsolved. 'Explainable AI reasoning' as a product feature is 2-4 years early. Build trust through empirical testing and human verification instead.

The good news for multi-language PMs: Claude operates in a language-independent conceptual space. Cross-language features are 2x+ higher in Claude 3.5 Haiku versus smaller models, scaling with model size. If you've validated AI behavior in English, larger models give higher confidence that behavior transfers across languages without per-language prompt engineering.

What to do

  1. Audit every product surface showing chain-of-thought or 'AI reasoning' to users — categorize by task difficulty. Add disclaimers or remove reasoning displays for high-difficulty categories where fabrication risk is highest

  2. Review all prompt templates for inadvertent 'hints' that could trigger motivated reasoning — flag any template passing expected answers or user-suggested conclusions into the model context

  3. Implement post-generation output filtering independent of model safety mechanisms — add a classification step after generation to catch content the model's safety features missed mid-sentence

  4. Redesign hallucination mitigation around 'recognition misfiring' — for RAG products, test whether injected context inflates the model's confidence for entities it doesn't actually know well

The bottom line

OpenAI just killed Sora after earning $2.1M on 3.3M downloads — torching a $1B Disney deal — proving that consumer AI without workflow retention is dead on arrival, while a New Mexico jury's $375M verdict against Meta established that your algorithms are product-liability targets that bypass Section 230, and Anthropic's research showed that every 'show your work' AI feature is fabricating reasoning on hard problems. The three takeaways: audit consumer AI features for Sora-pattern economics, add legal risk assessments to every PRD touching algorithmic content surfacing, and stop displaying chain-of-thought as evidence of AI reliability.