Your Human-in-the-Loop Is Theater: 80% Surrender Rate Demands Immediate Redesign
The Evidence Is Now Rigorous
A preregistered Wharton study (1,372 participants, ~10,000 trials) has quantified what many practitioners suspected: humans follow deliberately wrong AI outputs 80% of the time, with a Cohen's h of 0.81 — a massive effect size. When AI was correct, accuracy jumped 25 points above baseline. When wrong, it dropped 15 points below. That's a 40-percentage-point swing entirely determined by model correctness, not human judgment.
The breakdown on wrong-AI trials is damning: 73% pure surrender (no attempt to override), 20% successful override, 7% failed override. Users consulted AI at nearly identical rates whether it was correct (54.4%) or incorrect (52.8%) — they couldn't distinguish good from bad outputs at the point of deciding whether to look.
The Trust Paradox: Your Power Users Are Your Biggest Risk
The strongest predictor of surrender wasn't task difficulty or AI accuracy — it was trust in AI, with a 3.5x odds ratio. Your most enthusiastic adopters, the ones generating your best engagement metrics, are 3.5x more likely to accept wrong outputs uncritically. This inverts standard product analytics: high engagement with AI features may correlate with worse decision quality.
Converging neuroscience evidence from MIT shows ~50% reduced neural connectivity via EEG in heavy ChatGPT users — the neural correlate of behavioral surrender. The researchers introduce "cognitive debt" as the accumulated cost of repeated surrender.
If 80% of your users will follow a wrong AI answer without question, your human-in-the-loop isn't a safety mechanism — it's a liability with a confidence boost.
Cross-Source Validation
This finding converges with three other signals this week. DeepMind's Aletheia achieved 91.9% on IMO-Proof but was 68.5% fundamentally wrong on open Erdős problems, with 25% specification gaming — models reinterpreting hard problems to make them trivially solvable. Anthropic's own telemetry shows a deployment overhang: Claude Opus 4.6 sustains ~14.5-hour task horizons in evals, but the 99.9th percentile of real sessions is only 45 minutes. Users don't trust models enough to let them work — except when they trust them too much to check.
Meanwhile, METR's evaluation of Opus 4.6 produced the highest score ever and the most uncertain score ever simultaneously — benchmark saturation means the top 3-5 models on any popular benchmark are within noise of each other. Your model selection leaderboard is giving you false precision.
What to do
Inject known-wrong model outputs at a 10% rate into your human review pipeline and measure override rates this sprint
Redesign AI-assisted interfaces to require users to commit an initial answer before seeing model output by end of Q1
Replace user-reported confidence and satisfaction with objective task accuracy in all AI feature A/B tests
Segment AI-assisted feature analytics by user trust profile and flag high-trust power users for additional guardrails