The 39-Point Lie: Why Your AI Productivity Numbers Are Systematically Wrong
Four studies on AI productivity measurement
This quarter's four independent studies converge rather than contradict. Together they form the clearest indictment yet of how enterprises measure AI impact. The findings cut at the measurement layer, not the technology:
- METR Developer Study: Experienced open-source developers ran 19% slower with AI tools while believing they were 20% faster. That is a 39-percentage-point gap between sentiment and output.
- Brynjolfsson's 5,179-agent study: Junior customer support agents gained 34% productivity. Veterans saw near-zero improvement.
- Harvard/BCG Consultant Experiment: Large gains inside the competence boundary, 19% worse performance outside it.
- MIT NANDA (300 deployments, 150 executive interviews): 95% of enterprise GenAI pilots deliver no measurable P&L impact.
AI compresses rather than amplifies. It raises floors, not ceilings. If engineering leadership reports that AI tools are 'working great' based on team sentiment, the bill arrives twice: once in dollars, once in false confidence.
Where the Money Is vs. Where the Money Went
MIT NANDA locates pilot failure in organizational rigidity, not model quality. Firms bolt AI onto workflows they refuse to change. The specifics matter:
| Approach | Success Rate | Why |
|---|---|---|
| Vendor-purchased solutions | ~67% | Forces process adaptation |
| Internal builds | ~22% | Inherits existing dysfunction |
| Sales/marketing AI (most budget) | Low measurable ROI | Demos well, hard to attribute |
| Back-office automation (least budget) | Highest measurable ROI | Dull to present, clear to measure |
The Expert Pipeline Crisis Nobody Is Pricing In
Stanford's Canaries dashboard, built on ADP payroll data covering 1-in-6 US workers, shows employment for 22-to-25-year-olds falling in AI-exposed roles while rising for less-exposed peers. The chain is causal: AI helps juniors most, so junior roles become the most automatable, and automating them removes the apprenticeship rung that produces future experts. This is a leveraged bet that AI capability improves faster than the expertise reservoir drains. If it doesn't, the capability loss doesn't reverse on any normal timeline.
The Tradeoff Nobody Wants to Name
The board-deck version says deploy AI across the org and capture the gains. The complete version is harder. Capturing gains means moving money away from the function that asked loudest, sales and marketing, toward the function that didn't, back-office. It means retiring self-reported productivity metrics for instrumented A/B measurement. It means 70% of transformation budget going to process change and 30% to technology, the inverse of current allocation.
What to do
Kill all self-reported AI productivity metrics as primary indicators — implement instrumented A/B output measurement for every AI-assisted workflow by end of Q3
Rebalance AI investment portfolio: audit current split between customer-facing vs. back-office automation and shift 40% of customer-facing budget to back-office within two quarters
Design explicit expertise-formation pathways that coexist with AI augmentation — protect the apprenticeship rung for 22-25 year old hires
Mandate vendor-purchased solutions over internal builds for non-core AI use cases