Product & Strategy

The Product Desk

The Signal

Enterprise AI is stuck in a massive conversion crisis: 68% of 1

Novo Nordisk just showed the way through — they killed an expensive Anthropic-powered research tool that didn't deliver, redirected to process-automation agents that save $10–100M per week on clinical trials, and their CDO's mantra is 'if I can do it better in Excel

In Play

  1. Enterprise AI's 68/12 Conversion Crisis

    68% of S&P 500 AI partnerships remain pilots; only 12% are production relationships. Novo Nordisk killed exploratory AI and pivoted to agents saving $10-100M/week. Harvey hit $11B proving vertical AI beats horizontal bolting. Your enterprise buyers demand ROI against Excel, not demos.

    Ask Clarity
  2. Google's 2029 Post-Quantum Deadline Compresses Your Security Roadmap

    Google accelerated its PQC migration from NIST's 2035 to 2029, citing faster-than-expected quantum advances. Android 17 beta already ships PQC keys. 'Harvest now, decrypt later' attacks are active today. The White House may follow with a 2030 federal deadline — enterprise RFPs will require PQC readiness within 18 months.

    Ask Clarity
  3. Autonomous Agent Reality Check: All Models Score <1% Where Humans Score 100%

    ARC-AGI-3 benchmark: Gemini Pro 0.37%, GPT-5.4 0.26%, Opus 4.6 0.25%, Grok-4.20 0% — on tasks every human solves on first try. A vibe-coded app pentest found critical LFI and IDOR vulnerabilities. 48% of AI code contains hallucinated bugs. Ship human-in-the-loop, not autonomous workflows.

    Ask Clarity
  4. TurboQuant's 6x Compression Reshapes AI Feature Economics

    Google's TurboQuant delivers 6x KV-cache compression and 8x attention speedup on H100s with zero accuracy loss — no new hardware needed. AI memory stocks dropped 3-5%. Presenting at ICLR 2026 in April. Features shelved for inference cost are now viable. Simultaneously, VC consensus bet is AI infrastructure (Together AI at $7.5B), signaling compute costs will decline further.

    Ask Clarity
  5. Content Authenticity & Platform Trust in Structural Decline

    NYT, WSJ, and WaPo flagged for undisclosed AI use across thousands of articles. AI slop hit 3.3M TikTok followers in weeks. $8M streaming fraud from one actor went undetected by platforms. The 'AI;DR' backlash signals market bifurcation into 'verified human' premium and AI-commodity tiers. AI detection tools are fundamentally unreliable — invest in provenance, not detection.

    Ask Clarity

Deep Dives

Enterprise AI's Conversion Problem: From Pilot Purgatory to Production Revenue

The 68/12 Gap Is Your Biggest Strategic Signal This Quarter

New data from CB Insights reveals the starkest picture yet of where enterprise AI actually stands: across 1,000+ S&P 500 AI partnerships, 68% remain integrations, pilots, or co-marketing arrangements. Only 12% have matured into production vendor relationships. This number grew just 23% over two years — meaning the vast majority of enterprise AI spending is still experimental. If you're selling AI-powered products to enterprises, this reframes your entire GTM: the bottleneck isn't demand (budgets are allocated) or capability (models are good enough). It's the conversion from experiment to production value.

Novo Nordisk: The Case Study Every PM Needs

Novo Nordisk's CDO Stephanie Bova just provided the most instructive enterprise AI pivot of 2026. She killed Found Data — an expensive Anthropic Claude-powered tool that let researchers mine decades of clinical trial data for hidden trends. It sounded like a dream use case: big data, pattern recognition, potential drug discoveries. It was expensive to run and didn't lead to measurable advances. This is the canonical 'exploratory AI' failure mode.

Then she redirected to process-automation agents that detect clinical trial risks, auto-notify team leads via Microsoft Teams, and suggest remediation steps. The difference? Each week saved on a clinical trial is worth $10M–$100M in faster time-to-market. In Bova's words: 'If I can do it better and cheaper and more reliably in Excel, I'm going to tell you to stay in Excel.'

The 2026 enterprise buyer has moved from 'we need AI' to 'prove AI beats the baseline.' If you can't show measurably better outcomes than a spreadsheet, you're building Found Data.

Critically, Novo is running a multi-model orchestration architecture via Celonis — routing different agent tasks to Anthropic, OpenAI, or other providers based on task requirements. This validates model-agnostic architecture as a real enterprise pattern, not a theoretical best practice. Celonis ($1.6B in funding) is positioning as both the process mining data substrate and the model router — a platform play that PMs building enterprise AI need to either partner with or compete against.

Vertical AI Is Winning. Horizontal Bolting Is Losing.

The valuation data reinforces the lesson. Harvey (legal AI) hit $11B — a 3.5x jump in roughly one year — with $1B+ total funding. Granola (meeting transcription) reached $1.5B in a category many called crowded. Periodic Labs (AI for materials discovery) reached ~$7B after existing for just one year. Meanwhile, a Wall Street Journal report finds enterprises are using AI to build internal apps that chip away at seat-based SaaS revenue — not replacing Salesforce/SAP/Workday outright, but pressuring vendors on price with '80% good enough' tools built in days.

Superhuman's PMF Framework: The Conversion Methodology

Superhuman founder Rahul Vohra's extended PMF framework offers a concrete methodology for solving the conversion problem. When Superhuman scored just 22% on the Sean Ellis 'very disappointed' metric, Vohra didn't build new features. He narrowed the target persona — dropping non-core users to focus on VCs, CEOs, and founders — and the score jumped to 32% with zero product changes. A 45% improvement from spreadsheet work, not engineering. Systematic iteration on what core users loved then took it to 58%. His three added survey questions — 'Who is it best for?', 'What's the main benefit?', 'How can we improve?' — create a dual-track roadmap: amplify what 'very disappointed' users love, fix complaints from 'somewhat disappointed' users.

The highest-leverage product decision isn't what to build — it's who to build for first. Persona narrowing is cheaper and faster than feature work.

What to do

  1. Audit your enterprise sales funnel for pilot-to-production conversion rate and benchmark against the 12% industry baseline this sprint

  2. Categorize every AI initiative on your roadmap as either (a) exploratory/insight-discovery or (b) process-automation-with-measurable-ROI by end of week

  3. Run Vohra's extended PMF survey (Ellis question + persona, benefit, and improvement questions) segmented by user persona before looking at aggregate scores this quarter

  4. Evaluate Celonis or similar model orchestration layer for multi-provider AI routing before committing to a single LLM vendor

Google's 2029 Post-Quantum Deadline: Your Encryption Roadmap Just Lost 6 Years

The Timeline Shift

Google publicly accelerated its internal post-quantum cryptography migration from NIST's 2035 federal baseline to 2029 — a six-year compression driven by internal assessments that quantum hardware, error correction, and factoring algorithms are advancing faster than the industry assumed. This isn't a research preview: Android 17 beta is already shipping PQC key support, and NIST-vetted algorithms are ready for deployment at scale. The White House is reportedly considering pulling the federal deadline to 2030 or sooner, partly driven by anxiety over Chinese quantum breakthroughs.

Why This Matters for Your Product Now — Not in 2029

The 'harvest now, decrypt later' threat model is the key to understanding urgency. Google explicitly warned that adversaries are already recording encrypted traffic to decrypt with future quantum computers. Any data your product encrypts today with RSA or ECC that must remain confidential beyond 2030 is at risk right now. This applies to financial records, health data, enterprise IP, and any government-adjacent contracts.

Enterprise RFPs will start requiring PQC readiness within 18 months. Being able to answer 'yes, we've started migration' versus 'it's on our radar' is the difference between winning and losing deals.

Google's playbook mirrors their HTTPS push — Chrome didn't wait for regulation to flag HTTP as insecure. They set the de facto standard and let market pressure do the rest. There is no private sector mandate for PQC migration, but Google explicitly hopes its aggressive timeline pressures other companies to follow. Expect AWS, Azure, and Cloudflare to announce matching timelines within quarters — no hyperscaler can afford to be perceived as less secure.

The Architecture Requirement: Crypto-Agility

The practical implication isn't migrating everything this quarter. It's ensuring crypto-agility in your architecture — the ability to swap cryptographic algorithms without rewriting your application. Five sources converge on the same checklist:

  • Inventory every cryptographic dependency: TLS versions, key exchange algorithms, encryption-at-rest schemes, signing mechanisms
  • Flag which are quantum-vulnerable (RSA, ECC, most current key exchange)
  • Assess migration surface area and identify the longest-lead dependencies
  • Review data retention policies through the 'harvest now, decrypt later' lens

If your product handles data for government customers, the window is even tighter. Federal procurement officers will add PQC requirements to RFPs well before any formal mandate, following Google's narrative gravity.

What to do

  1. Commission a crypto-agility audit: have engineering inventory every cryptographic dependency and flag quantum-vulnerable algorithms by end of Q2

  2. Review data retention and encryption-at-rest policies specifically for data that must remain confidential past 2030

  3. Add 'post-quantum encryption readiness' to your competitive positioning matrix and monitor AWS/Azure/Cloudflare for matching announcements

  4. If selling to government: add a PQC compliance section to your federal sales deck this quarter

ARC-AGI-3 + Vibe Coding Pentests: The Dual Reality Check on AI Autonomy

The Benchmark That Should Recalibrate Every Agent Roadmap

ARC-AGI-3 launched with 135 mini games and ~1,000 levels, all designed to test fluid intelligence and agentic reasoning. The results are a cold shower: Gemini Pro scores 0.37%, GPT-5.4 High 0.26%, Opus 4.6 0.25%, and Grok-4.20 literally 0%. Meanwhile, 100% of untrained humans solve these tasks on first contact. The tasks require discovering rules, forming goals, and planning strategies with zero instructions — exactly the kind of autonomous reasoning that ambitious product roadmaps implicitly assume.

One crucial caveat: labs pushed ARC-AGI-2 from 3% to ~50% in under a year by spending millions on training. But as Mike Knoop notes, frontier labs are paying far more attention to V3 than earlier versions — meaning the benchmark-gaming cycle will repeat. Don't confuse benchmark score improvement with genuine reasoning capability improvement.

Vibe-Coded Apps Fail Security Review — Systematically

Independently, a grey-box pentest of a web app built 100% with Claude Opus 4.6 delivered damning results:

  • Critical LFI via unfiltered full_path parameter exposing /etc/passwd — path to remote code execution
  • IDOR on /employee/{guid} leaking emails, roles, and password hashes by harvesting GUIDs from a public API
  • Frontend running Vite 5.4.10 with three known CVEs

The pattern is structural: AI-generated code consistently skips input validation, enforces weak access controls, and ignores dependency management. Separately, data shows a 48% hallucination rate in AI code generation (o4-mini), and there's been a 1,300% spike in 'excessive agency' concerns among enterprise security teams.

AI-generated code optimizes for functional correctness, not defensive programming. If your velocity relies on AI code generation, budget 10-15% additional time for security review. That's still a massive net gain — but it's not free.

Software Quality Is Declining — And the Market Is Responding

Mario, founder of open-source agent Pi, reports that software quality is declining as more companies rely on agents. This isn't Luddism — it's a measurement observation from inside the ecosystem. AI coding agents compound errors without learning, generate excessive local complexity through individually optimal but collectively incoherent decisions, and suffer from low recall producing brittle codebases at scale. Tools like Expect (407K launch views) for agent browser testing and dev-browser (463K views) are emerging precisely because developers feel this quality gap viscerally.

The Positioning Opportunity

These findings create a nuanced competitive position. Your competitors using AI to ship faster are also accumulating more technical debt and security vulnerabilities. David Cramer (Sentry CEO) argues LLMs 'remove the barrier to get started, but create increasingly complex software which does not appear to be maintainable' and 'slow down long term velocity.' For PMs, the play is clear: your product's reliability, maintainability, and security posture are differentiators against AI-speed competitors who are building on a foundation of brittle code. Track 90-day code churn and maintenance burden on all AI-assisted features.

What to do

  1. Brief stakeholders on ARC-AGI-3 results using specific numbers (0.37% vs 100%) in your next roadmap review to recalibrate 'autonomous reasoning' feature expectations

  2. Add a mandatory security review gate for all AI-generated code before production — include input validation checks, access control review, and dependency auditing for any feature where >50% of code was AI-generated

  3. Add 'maintenance burden' and '90-day code churn' as tracked metrics for AI-assisted development initiatives starting this sprint

  4. Downgrade any 'fully autonomous agent' features from committed to experimental; redesign as human-in-the-loop with autonomous fast-paths for high-confidence tasks

The bottom line

Enterprise AI is in pilot purgatory — 68% of S&P 500 AI partnerships remain experimental, models score under 1% on tasks every human solves, and vibe-coded apps ship with critical security holes. But inference costs are about to drop 6x (Google TurboQuant), vertical AI companies that solve one domain deeply are hitting $7–11B valuations, and the PMs winning the conversion race (Novo Nordisk, Superhuman) aren't building more AI features — they're ruthlessly narrowing their persona, killing exploratory tools that can't beat Excel, and shipping process automation with measurable dollar-per-week ROI. Meanwhile, start your post-quantum crypto audit: Google says 2029 breaks everything, and your enterprise buyers will ask about it before you're ready.