Your AI Feature Needs Two Numbers Before Your Next Review
Three audiences are auditing AI claims — public markets, the press, and your own product data — and all three ask what the feature actually converted.
Why the market picked a seat count
An analyst opened Microsoft's earnings materials this week and reached for one number out of dozens: Microsoft 365 Copilot subscriber counts. Not because seats measure AI value well. Because a paid seat is the only AI figure an outsider can verify: a buyer chose to pay, per user, again. Usage minutes, prompt volume and "AI-assisted workflows" are constructions the vendor controls, so the market discounts them to zero. The Information's figures are pre-report; the earnings prints are the confirmation.
That question travels downward faster than most product teams expect. It lands in a roadmap review as "what did the AI feature convert?" — and an engagement chart is not an answer to it.
The same audit, pointed at product claims
Two examples from very different places share one structure. The feature was marketed on the easy metric while the hard one stayed unowned. The Bear Cave surfaced Forbes reporting that Axon's AI-generated police reports get facts wrong in public records, while the product is sold on time saved. Time saved is real. It is also not the gate. Once generative output enters a system of record, the binding number is the factual error rate, and nobody had published one. Tesla's Robotaxi line is the metric version of the same failure: "scaling," against operating miles that fell from 1.05M in Q1 to 0.75M in Q2, roughly a 30% decline.
An efficiency claim that has never been paired with an error rate is not a product metric. It is an unpriced promise.
| Claim being made | Easy metric used | Hard metric nobody owned |
|---|---|---|
| Enterprise AI is working | AI capex, ambition | Paid seats, retention |
| AI drafting saves time | Hours saved per user | Error rate, review policy |
| Autonomous fleet is scaling | The word "scaling" | Quarter-over-quarter miles |
Where the sources disagree
The Information frames AI capital as under-returning. TheSequence shows the opposite one layer down: Google Cloud growing 82% to $24.8B on enterprise AI demand, Alphabet planning roughly $180–190B of capex, and private marks re-rating hard, with Databricks moving from $134B to $188B in five months. Both readings hold. The reconciliation is the useful part. Infrastructure demand is genuine; per-feature paid conversion is what remains unproven. Apple is the uncomfortable control group, up 23% year to date on a real iPhone upgrade cycle with minimal AI spend. Having no AI story was not punished.
The move
Build the internal version of the seat metric before someone hands you theirs. Four numbers per AI feature: paid attach rate, retention delta between feature users and non-users, cost per successful outcome, and a documented error rate with its review policy. Most teams can assemble the first three from existing telemetry. Almost none have the fourth, which is exactly why it is the one that decides the launch review.
What to do
Instrument paid attach and 30-day retention for every shipped AI feature this sprint, and report cost per successful outcome next to usage.
Define a documented error-rate threshold and human-review policy for any AI output that enters a system of record, before the next launch review.
Pull the raw quarter-over-quarter trend behind any AI growth metric in your next exec deck before it ships.