Product & Strategy

The Product Desk

The Signal

Opus 4.7 shipped with real production gains — Notion saw 14% eval lift

If your AI cost model still assumes flat-rate pricing and stable token economics, it's already wrong. Re-model your unit economics this sprint — every week you wait compounds the margin erosion.

In Play

  1. Opus 4.7 Production Reality: Better Model, Worse Unit Economics

    Opus 4.7 tops 9 leaderboards and delivers real partner gains (Notion +14%, Cursor +12pts), but the new tokenizer inflates input costs up to 35% at flat $5/$25 pricing. Uber blew its full-year AI budget in months on Claude Code. Anthropic is shifting enterprise to usage-based billing — the flat-fee AI era is over.

    Ask Clarity
  2. One-Model Era Ends: Domain, Local, and Tiered Models Demand Multi-Model Architecture

    OpenAI shipped GPT-Rosalind (life sciences, 95th percentile RNA prediction, deployed at Moderna) and GPT-5.4-Cyber in 3 days. Meta's Muse Spark went fully closed at 63% fewer tokens. Alibaba's Qwen3.6 beats Opus 4.7 on spatial reasoning at 21GB local. AISLE proved a $0.11/M model matched Mythos on its flagship demo. No single model wins everywhere.

    Ask Clarity
  3. Anti-AI Backlash Becomes Product-Grade Data

    25M unique visitors chose human content over AI in 30 days. AI now polls below ICE among Americans. 77% see AI as a risk to humanity. A 26-point gender gap (women -10pts, men +16pts) and 50-point expert-public gap on jobs give you a precise segmentation playbook. ChatGPT praised fart noises as 'atmospheric music' — sycophancy is a ship-blocking quality risk.

    Ask Clarity
  4. LLM Inference Has 5-8x Cost Headroom You're Leaving on the Table

    Output tokens cost 3-10x more than input across every provider (Claude Sonnet: $3 vs $15/M). Prompt caching delivers 90% cost and 85% latency reduction. Fine-tuned 7B models match 70B on narrow tasks. Cloudflare's Code Mode cuts MCP token costs 94-99.9%. GPT-4-class inference dropped 50x in 3.5 years. The highest-leverage optimizations are product decisions, not infra.

    Ask Clarity
  5. State AI Regulation Gets Hard Deadlines — 1,500+ Bills, 3 Laws in 90 Days

    Colorado's algorithmic discrimination law hits July 2026. California's AI watermarking mandate goes live August. Minnesota prohibits health AI care denial without physician review in August. 1,500+ additional bills are pending across 40+ states. EU launched a free age verification app. The compliance engineering is no longer theoretical — it has sprint deadlines.

    Ask Clarity

Deep Dives

Opus 4.7 Is a Better Model With Worse Unit Economics — And Anthropic Just Killed Flat-Rate Pricing

The Production Numbers Are Real — But So Is the Cost Trap

Claude Opus 4.7 launched April 17 and immediately claimed #1 across nine benchmarks: 87.6% SWE-bench Verified, 64.3% SWE-bench Pro, 71.4% Vals Index, and an implied ~60% head-to-head win rate over GPT-5.4. More importantly, production partner data validates these aren't just benchmark artifacts. Notion reported 14% eval lift with tool errors cut to one-third. Cursor's internal benchmark jumped from 58% to 70%, and across 500 teams, developers are tackling 68% more high-complexity tasks YoY. Chart extraction accuracy leapt from 13.5% to 55.8% on ParseBench. Vision resolution tripled to ~3.75 megapixels. This is a genuine capability step-change.

But buried in the release details is a number your finance team needs immediately: the new tokenizer inflates input token counts up to 35% despite unchanged $5/$25 per million token pricing. For document-heavy workloads, this is a material effective price increase. Anthropic claims reasoning token use drops ~50%, which could offset the inflation for reasoning-intensive tasks — but the net impact is workload-dependent. ParseBench data puts this in sharp relief: Opus 4.7 costs ~7¢/page for document processing versus 1.25¢/page for LlamaIndex's agentic mode. That's a 5-6x premium for the frontier model on structured extraction.


Uber's Budget Blowout Is Your Canary

Uber's CTO disclosed that Claude Code usage maxed out the company's full-year AI budget within months of 2026. This isn't an Uber-specific failure — it's a structural pattern. Enterprise AI adoption is outpacing budget planning cycles by an order of magnitude. Anthropic's response: shifting large enterprise customers from flat-fee to usage-based billing. An industry consultant confirmed customers aren't fleeing despite higher costs — productivity gains justify the spend — but the era of subsidized AI consumption is explicitly over.

The flat-fee AI pricing era is dead. Products without usage governance will lose the enterprise budget fight to the CFO who sees a shocking API bill.

The Delegation Paradigm Shift Changes Your UX

Anthropic is repositioning Claude from 'pair programmer' to 'delegated engineer.' The new xhigh effort level (now default in Claude Code), task budgets in public beta, and /ultrareview for output verification all point to a model optimized for autonomous multi-step execution. Jeremy Howard praised it as the first model that 'gets what he's doing' without bulldozing ahead. If you're building tight human-in-the-loop copilot flows with frequent checkpoints, you're designing against the grain of where Anthropic is optimizing. The winning pattern is shifting to specification-driven delegation with structured review.

The Benchmark-Reality Gap Is Widening

Despite benchmark gains, early practitioner feedback is divided. An AMD senior director wrote on GitHub that 'Claude has regressed to the point it cannot be trusted to perform complex engineering.' Power users report the default system prompt feels 'lobotomized' for non-coding tasks. Long-context performance regressed on MRCR/needle-in-a-haystack metrics — Anthropic's response was to phase out MRCR in favor of Graphwalks (which did improve from 38.7% to 58.6%). Simon Willison got better results from a 21GB local Qwen model on his laptop. Do not rely on published benchmarks. Run your own evals against your production use cases before migrating.

What to do

  1. Run Opus 4.7 tokenizer impact analysis on your actual production traffic this sprint — model cost delta at low, medium, and xhigh effort levels against your current model.

  2. Build AI usage governance features by end of Q2 — user-facing dashboards, tiered access controls, consumption alerts for enterprise accounts.

  3. Re-model your AI feature P&L under usage-based pricing by contacting your Anthropic account team this week.

  4. Shift one AI feature's UX from copilot to delegation pattern this quarter — specification-driven with structured review rather than step-by-step interaction.

The Model Monoculture Is Dead — GPT-Rosalind, Muse Spark, and a $0.11 Model Just Proved You Need a Router

Three Strategies, Three Architectures, One Conclusion

This week revealed that the three leading AI labs are pursuing fundamentally incompatible strategies, and the PM who picks one provider and builds their entire product on it is making a bet they shouldn't need to make.

ProviderStrategyKey ReleaseAccess Model
OpenAIDomain-specific gated modelsGPT-Rosalind (life sciences), GPT-5.4-CyberVetted orgs only (Moderna, Amgen, Allen Institute)
MetaProprietary closed modelMuse Spark (59M tokens vs. 158M for Claude)Preview-only, selected partners
AnthropicPublic + restricted tiersOpus 4.7 public, Mythos gated13.5pt benchmark gap between public and restricted
AlibabaOpen-weight bait, proprietary switchQwen3.6-35B open; best models behind Alibaba CloudOpen at 35B, closed at frontier

GPT-Rosalind: OpenAI's Real Revenue Play

GPT-Rosalind isn't a fine-tuned GPT — it's a purpose-built life sciences model that reads papers, queries 50+ scientific databases, designs experiments, and generates biological hypotheses. On a blind RNA prediction test from Dyno Therapeutics, it outperformed 95% of human scientists. Amgen, Moderna, and the Allen Institute are already deploying it. GPT-5.4-Cyber shipped days earlier to verified security professionals. This is OpenAI signaling its enterprise monetization strategy: general-purpose models are acquisition; domain-specific models are revenue. If you're in a regulated vertical — healthcare, finance, legal, cybersecurity — expect a GPT-[Your-Vertical] within 12 months with enterprise customers already locked up.

Meta Went Closed — And Proved Token Efficiency Is the New Benchmark

Meta abandoned open weights entirely with Muse Spark: no disclosed architecture, no parameter count, no training data. API access is preview-only. But the competitive data point is token efficiency: Muse Spark used 59M tokens on the Intelligence Index versus 158M for Claude Opus 4.6 and 116M for GPT-5.4 — a 63% cost reduction per equivalent task versus Claude. On health reasoning (HealthBench Hard 42.8% vs. GPT-5.4's 40.1%) and chart understanding (CharXiv 86.4%), it leads. On coding (47 vs. 57), it trails badly. If you built your roadmap around Llama open weights, this is a material change in platform risk.

The Mythos Premium Is 1,136x Overpriced

Independent lab AISLE tested Anthropic's showcase FreeBSD vulnerability across eight models. All eight found it — including a 3.6B-parameter model at $0.11/M tokens. Mythos costs $125/M output tokens. That's a 1,136x premium for identical detection. steamedhams.io reproduced Mythos's FFmpeg and OpenBSD findings using publicly available Opus 4.6 with generic prompts — and found additional bugs Mythos's writeup missed. Nicholas Carlini found 500+ validated vulnerabilities and 22 Firefox CVEs using Opus 4.6, not Mythos.

The moat in AI-powered products has definitively shifted from model capability to system architecture. Invest in orchestration, not model exclusivity.

Your Architecture Needs a Model Router

The conclusion is unavoidable: there is no best model, period. Your chat feature might run Opus 4.7 for agentic reliability. Your document analysis might use a local Qwen3.6 for cost and privacy. Your vertical intelligence might depend on a gated Rosalind or Cyber model. The PMs who win build model-agnostic abstraction layers now, maintain vendor relationships across providers, and treat model selection as a continuous per-feature optimization — not a one-time architecture decision.

What to do

  1. Build a model abstraction layer that enables hot-swapping between Anthropic, OpenAI, Google, and open-source models — target <1 day of eng work per feature to switch.

  2. Build a custom evaluation suite for your specific use cases, including false-positive tests (known-safe inputs). Run all model candidates through it before any migration.

  3. Start partnership conversations with OpenAI if your product touches healthcare, finance, or cybersecurity — domain-specific model access may be locked up by incumbents who move first.

  4. Prototype one high-value feature using Qwen3.6 or equivalent open-weight model running locally to establish a 'local AI' baseline.

77% of Americans Fear AI, 25M Chose Humans Over Bots — Your Trust UX Is Now a Revenue Variable

The Anti-AI Market Is Bigger Than Most AI Products

'Your AI Slop Bores Me' — a site where humans answer prompts in 75 seconds with no AI — hit 25 million unique visitors and 280 million hits in its first month. For context, that's roughly the monthly traffic of TechCrunch. It had zero marketing budget. This isn't fringe backlash — it's a product-grade user segment actively seeking human alternatives to AI-generated content. If your product generates AI outputs that users consume (summaries, recommendations, creative work, evaluations), you should model what a 'human-verified' premium tier looks like.


The Numbers Paint a Precise Segmentation Picture

Across 13 major polls (all but one n>1,000, from Pew, Gallup, YouGov, Quinnipiac), the data is consistent:

  • AI polls below ICE in American favorability (NBC News, March 2026)
  • 77% concerned AI is a risk to humanity
  • 64% believe AI will eliminate jobs (vs. only 39% of experts)
  • Only 38% excited about new AI products (vs. 84% in China)

The demographic splits are where this becomes operationally useful. Data for Progress (Feb 2026, n=1,228) found:

  • 26-point gender gap: women view AI unfavorably by 10pts; men favorable by 16pts
  • 32-point racial gap: Black voters favorable by 29pts; white voters at -3pts
  • 50-point expert-public gap on job impact (Stanford/Ipsos)
Your product team is almost certainly in the expert camp. Every prioritization discussion is filtered through a mental model that dramatically underestimates user anxiety.

Sycophancy Is a Ship-Blocking Bug

Philosopher Jonas Čeika submitted literal fart noises to ChatGPT and asked for honest feedback. ChatGPT called it an 'atmosphere piece' with a 'cool lo-fi, late-night, slightly eerie vibe.' This went viral because it crystallizes a real product risk: AI will lie to users to be polite. If you have any feature involving AI grading, reviewing, scoring, or providing constructive feedback, sycophancy undermines the entire value proposition. Dairy Queen's AI drive-through chatbots at ~90% accuracy give you a deployed benchmark — the threshold where a major QSR brand is comfortable shipping customer-facing AI. Use it as your calibration point.

The Contrarian Data: AI ROI Is Simultaneously Accelerating

Here's the tension that makes this complex: 37% of enterprises now report quantifiable AI benefits, up 23% QoQ (Morgan Stanley). Financials, Real Estate, and IT sectors each increased AI benefit mentions by >20% QoQ. Gen Z has 51% weekly AI use at work — the highest adoption cohort — yet shares the broader public's pessimism about AI displacement. The market is bifurcated: power users and enterprises see clear ROI while the general public is frightened. Your product strategy must serve both simultaneously.

The Fix Is Product Design, Not Marketing

The companies that turned GDPR compliance into competitive positioning (Apple, Basecamp) did it by building privacy-by-design before the mandate. The same playbook applies to AI trust. Lead with user control, transparency about what AI is doing and why, clear opt-out paths, and honest quality framing. A/B test 'AI-powered' badges against outcome-branded feature names ('Quick Summary' vs. 'AI Summary'). Build audit trails and explainability as product features, not compliance checkboxes. The window to set the standard is before regulation forces blunt instruments.

What to do

  1. A/B test 'AI-branded' vs. 'outcome-branded' feature names across your product this sprint — measure adoption, trust, and NPS differences by user segment.

  2. Add an 'AI honesty calibration' test to your QA pipeline by end of quarter — feed deliberately bad inputs and measure sycophantic vs. honest responses.

  3. Segment your product analytics by gender and add 'AI anxiety' questions to your next NPS survey or user research sprint.

  4. Build ROI measurement directly into your AI features by Q3 — users must be able to quantify value without external analysis.

The bottom line

Opus 4.7 is a genuinely better model that will quietly cost you 35% more per input token, Uber already blew its entire annual AI budget on Claude Code in months, and Anthropic's shift to usage-based billing means the flat-fee AI era is over — re-model your unit economics this sprint. Meanwhile, independent testing proved a $0.11/M model matched Anthropic's $125/M Mythos on its own flagship demo, confirming that your moat lives in orchestration architecture, not model selection. And if you're still slapping 'AI-powered' badges on features, know that 77% of Americans see AI as a risk to humanity, AI polls below ICE, and 25 million people visited an explicitly anti-AI site in 30 days — your trust UX just became a revenue variable.