Substitution Failed; Supervision Is Where the Work Went
Uber's agent program shows the same physics with the opposite outcome, and the difference is what each company funded before it switched the agents on.
The mechanic is a cost transfer, not a cost reduction
AI does not remove engineering labor from the system. It moves the labor from authoring to verifying, and verification is harder to hire for, harder to measure, and invisible in velocity dashboards until it turns up in an incident report. A business case that counted the first effect and not the second is overstated by construction, not by accident. The tell sits inside Meta's own numbers: the 10% reduction went ahead even after the agent program was scrapped, which means the structural cost pressure never depended on AI performance in the first place. Letting AI carry the narrative for a cut it did not earn buys accountability for productivity gains that may never be evidenced.
The counter-case, and why it is not a contradiction
Pivot 5's account of Uber is the most concrete public dataset on production agentic development: 2,500 registered skills at 20,000+ daily invocations, more than 1,000 tools behind a single MCP gateway, 100M+ daily LLM requests through a centralized gateway, over 70% of pull requests agent-authored, and code shipped per engineer doubled year over year. Read as a scoreboard, that is the opposite result. Read the restraint instead. Uber's coding agent stops at a draft pull request rather than hitting shared CI, because unvalidated agent output was burning pipeline capacity, and maintenance diffs are size-capped and batched to Sundays.
| Dimension | Meta's Project OT | Uber's agent platform |
|---|---|---|
| What was funded | Generation plus headcount removal | Generation plus gateway, skill registry, CI capacity |
| Throttle | None disclosed | Draft-PR halt, Sunday-batched diffs |
| Claimed outcome | Up to 60% smaller teams | Doubled output per retained engineer |
| What broke | Reliability, then the plan | Pipeline capacity, then rate-limited |
Agent throughput has already outrun validation infrastructure. The mature response was rate-limiting, not more generation.
The three-to-five-year bill
The a16z crypto essay in the source set names the second-order cost the P&L never sees. Judgment is accumulated rather than innate, produced by supervised correction, someone catching the same error repeatedly until it stops recurring. The function that survives inside an LLM workflow is precisely that one: verification, drift-catching, correction. Delete the entry-level layer and the drafting savings book this fiscal year while the only mechanism that produces the seniors who verify machine output is severed. The third path almost nobody has designed for is to redirect junior roles from drafting to supervised verification, which captures most of the savings and keeps the loop running. The discarded byproduct is the valuable one: a logged corpus of senior corrections is a proprietary record of how a firm actually exercises judgment, and no general-purpose model will ever ingest it.
Where the evidence is thin
Techpresso concedes its magnitudes are directional: a digest of documents Reuters viewed, with an internal discrepancy on the settlement figure in the same issue. Verify before citing externally. The honest skeptic's objection is that one program at one company proves little. The skeptic is right about the sample size. What the skeptic still has to explain is why the firm with the most compute, the deepest internal tooling and the highest talent density could not make agent-supervised engineering hold. The tradeoff is not adoption speed against caution. It is whether adoption gets priced with someone else's numbers or with numbers earned in-house.
What to do
Instrument change-failure rate, incident volume and share of engineering time spent on remediation, split by AI-assisted versus human-authored code paths, and put the split on the next operating review.
Gate every AI-attributed headcount reduction in the FY27 plan on two consecutive quarters of stable reliability metrics, and book structural cost actions separately from AI savings.
Fund validation capacity — CI throughput, review automation and a supervised-verification track for entry-level roles — as a named platform line item this planning cycle.