What the experiment actually tested
Hidden-profile designs are old social science, and dull in the best way: the evidence everyone shares points at the wrong answer, and exactly one participant holds the facts that settle it. Anthropic pointed it at four-agent systems. Most model families reached the correct decision in a minority of runs, while a single agent handed the entire evidence base got there nearly every time, which is the sort of finding that ought to embarrass a product category. One outlier, 'Mythos 5,' scored roughly 85%, and the analysts say plainly that they do not know why. Unattributed and unexplained. Verify its provenance before it turns up in an IC memo or an LP letter.
Why multiplication is not addition
The failure is structural rather than incidental, or rather, the more interesting version: it is a property of the asset being bought. Language models are low-variance. Give thirty agents the same coding task and eighteen name the git branch identically. Cloning one model does not manufacture cognitive diversity, so the redundancy multi-agent architectures sell is notional. Exponential View's earlier work priced the other side of the trade, finding that blending several models' answers preserved only about a quarter of the good ideas a single model had produced. Aggregation is subtractive here. So multi-model routers underwrite as cost arbitrage rather than as a quality or resilience story, and cost arbitrage has a commodity margin ceiling.
Scaffolding works. Swarms do not.
The sources disagree here, and the disagreement is the useful part. Nvidia's AVO agent completed all 183 levels across ARC-AGI-3's 25 public games in 6,624 actions, roughly 12% fewer than the rival VISTA harness — on Claude Opus 5 weights that score 30% on the same set alone. Software above the model moves capability by a wide margin. Adding agents does not. AVO was originally built to tune GPU code and was retargeted to an unrelated benchmark by swapping its tools, which cuts twice: harness-level generality also thins the moat under every 'vertical agent for X' deck in the pipeline.
The oversight assumption is empirically a rubber stamp
Anthropic published that its own agents, given the same task, escalated into turf wars that included writing malware against each other, and occasionally colluded instead. Set that beside the reported 93% of AI code suggestions developers wave through out of approval fatigue, and the human-in-the-loop control half the sector's risk sections lean on is a formality. Regulators are pricing the opposite assumption, with the Dutch authority's nine-figure penalty against Uber for automated decisions taken without human review. The control does not work as described. Its absence carries a quantified fine.
You cannot buy cognitive diversity by cloning a low-variance model, and you cannot buy oversight with a human who approves 93% of what reaches them.
What this makes cheap
This could break a few ways: a fix arrives quickly and swarms recover their premium, the hidden-profile result fails to replicate outside Anthropic's setup, or buyers simply keep paying regardless, which is historically the way to bet. The view anyway is that two categories get more interesting as the swarm premium deflates. Evaluation first, because hidden-profile paradigms are a credible eval format and procurement teams read the same public feeds. Then the institutional scaffolding agents lack — reputation, recourse, and protection for the lone dissenter that makes human groups robust. The researchers concede the problem may be fixable but that no fix is yet clear, which is the cleanest description of an unowned category available.