Enterprise AI's Paradox: 95% Fail, Budgets Still Growing — Finding the 5%
The Contradiction That Reprices Your Book
Four independent studies converged this week on an uncomfortable truth, and one corporate survey contradicts the bear case they imply — together creating the most actionable intelligence pattern in today's briefing.
MIT NANDA (300 deployments, 150 exec interviews, 350 employee surveys): 95% of enterprise GenAI pilots delivered no measurable P&L impact. Only 5% produced rapid revenue acceleration. The blocker isn't model quality — it's the 'learning gap': organizations bolting AI onto unchanged workflows.
RBC CIO Survey: enterprises are creating net-new AI budgets, not cannibalizing existing software spend. 50%+ already run AI in production; another 35% within six months. Token costs aren't slowing adoption.
The market is simultaneously proving AI budgets are expanding AND that 95% of those dollars generate zero return. The 5% that works is the only thing worth funding.
What Separates the 5% From the 95%
Three structural findings rewrite sourcing priorities:
| Finding | Data | Investment Filter |
|---|---|---|
| Buy beats build 3:1 | ~67% vendor success vs ~22% internal | Tailwind for vertical AI SaaS; headwind for 'enterprises self-assemble' infra |
| ROI hides in back-office | Sales/marketing AI: most budget, least return | Contrarian alpha in underfunded finance ops, compliance, document processing |
| Self-reported gains are fiction | METR RCT: devs 19% slower, believed 20% faster | Discount any pitch with sentiment-based KPIs |
The productivity data adds nuance: AI lifts novices +34% but veterans ~0% (Brynjolfsson, n=5,179). Stanford's Canaries dashboard — built on ADP data covering 1-in-6 US workers — already shows employment falling for 22-25-year-olds in AI-exposed roles.
The Coinbase Counter-Example
Against the 95% failure backdrop, Coinbase just proved the opposite extreme: it cut AI spend roughly 50% while increasing token usage by defaulting to Chinese open-weight models (GLM 5.2, Kimi 2.7). This isn't a pilot — it's a public company optimizing at the P&L level. The key: Coinbase didn't bolt AI onto an unchanged process. It restructured its inference economics as a deliberate operational move.
The pattern becomes clear: the 5% that works either forces workflow redesign (the MIT finding) or owns its inference cost structure (the Coinbase finding). Most of your deal flow does neither.
What This Means for Your Portfolio
The net-additive budget finding dismantles the 'AI just eats SaaS' bear thesis — TAM is expanding, not zero-sum. But 95% failure means most of those dollars will churn. The alpha is in identifying which companies force the workflow change that makes value stick.
What to do
Rebuild AI app-layer diligence around three gates: (a) does the product force workflow redesign, (b) vendor-deployed not customer-built, (c) inference contribution margin at scale
Audit existing portfolio for self-reported productivity KPIs and request controlled output metrics plus token-cost-per-unit trends from each AI company
Re-weight sourcing toward back-office automation AI (finance ops, compliance, document processing) and away from sales/marketing AI
Open a thesis on AI-agent validation/testing as a category — the 'testing the 95%' wedge