The 39-Point Productivity Lie: Four Independent Studies Prove Your AI Metrics Are Measuring Vibes
The Convergence
Four independent studies this cycle — 5,179 support agents, a 16-developer RCT, ~758 BCG consultants, and 300 enterprise deployments — converged on one read. AI productivity gains are real, heterogeneous, jagged, and systematically over-reported by the people experiencing them.
METR's randomized trial (arXiv 2507.09089) is the cleanest of the four. Experienced open-source developers working on their own repositories ran 19% slower with AI assistance while believing they were 20% faster, after expecting a 24% speed-up going in. That is a ~39-point gap between measured time and self-report, under randomization. The randomization is what moves it past anecdote.
The Distribution Matters More Than the Average
Brynjolfsson's field experiment reports a 14% average productivity lift. That average is the least useful number in the study. Decompose it and you get +34% for novices, near-zero for veterans. The mean tells you to roll out uniformly. The distribution tells you it is an onboarding accelerant, not a senior-IC multiplier.
Dell'Acqua's BCG study adds the second axis, the jagged frontier. Inside it, consultants did 12.2% more tasks, 25% faster, at higher quality. One task outside it, with no visible boundary, ran 19% less likely to be correct. The lift is conditional on the task, not just the user.
A pooled average across a mixed population is a number that describes no one on your team.
The Enterprise Reality Check
MIT NANDA's field study (150 interviews, 350 surveys, 300 deployments) carries the sobering figure: 95% of enterprise GenAI pilots delivered zero measurable P&L impact. The blocker was workflow integration, the 'learning gap,' not model quality. Vendor-bought solutions succeeded at ~67% against ~22% for internal builds. ROI concentrated in back-office automation, not the sales and marketing tools that consumed most of the budget.
Glean's 6,000-worker survey supplies the mechanism. AI saves time, and much of it goes back into cleanup and rework. A naive completion-rate metric reads high because it does not instrument that downstream cost.
| Study | n / Design | Key Effect | What Your Metric Misses |
|---|---|---|---|
| METR | 16 devs, 246 tasks, RCT | -19% speed, +39pt perception gap | Self-report is invalid |
| Brynjolfsson | 5,179 agents, field exp | +34% novice, ~0% veteran | Averages mask bimodality |
| BCG/Dell'Acqua | ~758, controlled exp | +25% in-frontier, -19% outside | Adjacent tasks yield opposite results |
| MIT NANDA | 300 deployments | 95% = zero P&L; vendor 67% vs build 22% | Workflow integration is the blocker |
What This Changes
Surveys, NPS, and 'felt faster' metrics carry a known upward bias that METR quantifies at 39 points. The thing those metrics don't tell you is actual completion time. The replacement is a measured holdout, stratified by skill tier and task-frontier alignment, instrumenting completion time and quality. The holdout costs little next to redirecting a team on pooled averages that were measuring noise.
What to do
Replace all self-reported AI productivity metrics with a randomized or matched-control holdout measuring task completion time and quality, stratified by experience tier
Decompose your AI tool's effect by skill tier: measure novice vs veteran separately and report both, never just the pooled average
Map your AI tool's 'jagged frontier' — catalog task categories where it helps vs. hurts and gate usage accordingly
Instrument cleanup/rework time downstream of AI-assisted work before reporting productivity gains