The AI Measurement Crisis: 39 Points of Self-Deception
The AI Success Metric Most Teams Trust Is the One Most Likely to Mislead Them
A developer finishes a task with AI assistance, closes the editor, and reports that it went faster. The clock says otherwise. Three studies published this cycle measure that gap, and the gap is the whole story. The data is uncomfortable because it contradicts what the people doing the work say about the work.
Developers were 19% slower with AI but believed they were 20% faster — a 39-point perception gap that invalidates any AI feature success measurement based on user sentiment.
METR ran a randomized trial with 16 experienced open-source developers across 246 real tasks on their own codebases. With AI, they were measurably slower. Surveyed afterward, they said the opposite. Before starting, they predicted a 24% speed-up. This is not a calibration error you can survey your way out of. The people closest to the work were the most wrong about it.
The Inverted Expertise Curve
Brynjolfsson's study of 5,179 customer support agents, published in QJE, maps the gain by skill level: +34% for novices, approximately zero for veterans. The BCG/Harvard study adds the part that should worry anyone shipping to experts. Inside AI's capability boundary, consultants did 12.2% more tasks 25% faster. Outside that boundary, AI-assisted consultants were 19% less likely to reach the correct answer than the control group.
Building AI features for power users because they are the loudest stakeholders targets the segment with the lowest measured impact. On the harder tasks it may degrade their work while they applaud the speed they did not actually gain.
95% Pilot Failure Has a Specific Cause
MIT NANDA surveyed 150 executives and 350 employees and reviewed 300 public AI deployments. The headline is that 95% of GenAI pilots delivered no P&L impact. The cause matters more than the number. Companies bolt AI onto workflows nobody redesigned and expect a different result. NANDA calls the missing piece the learning gap. That is the actual blocker.
Two numbers sharpen the build-versus-buy call: vendor-purchased AI solutions succeed ~67% of the time versus ~22% for internal builds. And the ROI that showed up came from back-office automation, not the customer-facing tools absorbing most of the budget.
The Cleanup Tax Is Real
Glean's survey of 6,000 digital workers reaches the same place from a different door. AI saves time, and much of that time goes back into cleanup. What teams report as productivity is gross, not net. Measure net output — code that shipped and stayed shipped, not code that was generated.
What This Means for the Next Sprint
A sprint review that logs "users love the AI feature" off survey data may be celebrating a feature that slows the work while manufacturing the feeling of speed. The fix is behavioral instrumentation, not sentiment: actual task completion time, error rate, rework rate, and net output quality after corrections, each measured independently of what users believe happened.
What to do
Replace all self-reported AI satisfaction metrics with behavioral instrumentation (task completion time, error rate, rework rate) by end of next sprint
Segment your AI feature usage data by user expertise level and run a novice-vs-expert impact analysis this sprint
Present MIT NANDA data (67% vendor success vs 22% internal build) at next build-vs-buy decision point
Reposition AI features in product narrative as 'novice accelerators' and 'skill-gap closers' rather than expert productivity tools