The two different jobs people call "the harness"
Scaffolding does two structurally different things, and these results only look contradictory until they are separated. Adding steps whose output a verifier can check compounds capability: search and backtracking policy, retained reasoning, compaction, verify-before-commit. Adding voices that a reducer must reconcile averages capability away. The first is engineering. The second is a vote.
The reliability arithmetic explains why the first one works at all. A loop adds no capability. It amplifies whatever per-step reliability already exists, and below a threshold it amplifies errors instead.
| Per-step reliability | 10 steps | 20 steps | 50 steps |
|---|
| 0.95 | ~60% | ~36% | ~8% |
| 0.98 | ~82% | ~67% | ~36% |
| 0.99 | ~90% | ~82% | ~61% |
| 0.995 | ~95% | ~90% | ~78% |
Moving from 0.95 to 0.99 per step is a 2.3x improvement in 20-step success. That is where the engineering hours belong. Verification is cheaper than reliability: a passing test suite, a type check, or a fresh-context retry raises effective per-step reliability faster than any amount of prompt tuning. Derive a maximum unattended horizon from the measured number and force a checkpoint there. Same discipline as a retry budget on a downstream call.
Why the redundancy is not redundant
Given one coding task, 18 of 30 agents chose an identical git branch name, per Exponential View. That is 60% agreement on free text with effectively unbounded entropy, which measures mode collapse rather than coincidence. Every reliability pattern borrowed from distributed systems assumes independent failure domains: N-of-M voting, retry with a different seed, LLM-as-judge, ensemble guardrails. Same-family replicas do not supply that. Correlated errors do not fail noisily. They return unanimous, confident, wrong answers, the hardest incident class to detect. Move critical judging to a different vendor lineage than the generator, and track cross-vendor disagreement rate as a live signal. It doubles as a drift detector.
The refactor that keeps the legitimate reason to fan out
Convert multi-agent from a voting system into a retrieval system. Agents return structured claims carrying evidence IDs plus a mandatory "facts only I hold" field. A deterministic reducer unions the evidence with no model judgment at reduce time. One agent then decides over the full union. Instrument evidence coverage rate, the fraction of unique facts that actually reached the decision step, because that metric is the production canary for the hidden-profile failure. Prevent anchoring mechanically with sealed-bid submission. Instructing an agent to consider dissenting views does not repair a low-variance prior. Withholding the consensus does.
Which components to write for deletion
| Component | Absorption status | How to architect it now |
|---|
| Tool calling | Absorbed via in-environment RL | Use native; delete custom prompt scaffolding |
| Context compaction | Absorbed natively in newest coding models | Adapter behind a capability flag |
| Retained reasoning | Absorbed post-o1 | Configure, do not reimplement |
| Tool selection, memory, orchestration | Next in line | Thin, swappable, no business logic coupling |
| Permissions, identity, audit | Absorption-proof: human-facing | Deterministic, external, versioned |
Anthropic deleted 80% of Claude Code's system prompt without losing capability, per Latent.Space. Good craftsmanship, and it makes "how much harness can we delete at constant score" a real progress metric, provided the deletions land in tranches and a separate adversarial suite stays in place, because partially absorbed guardrails fail on the tail an aggregate benchmark never samples.
What not to cite in a design review
The 106-task spread arrives with no named owner and no published methodology. Nvidia's perfect run was on the public set, with a harness retargeted to that interface and no token cost or wall-clock published. Directionally believable, quantitatively uncitable. Build a frozen suite in-house and cite that. Two measurement rules make it valid: freeze the harness to compare models, freeze the model to compare harnesses, and report actions-per-solved-task and tokens-per-solved-task next to pass rate. Measured token price elasticity of 1.2-1.8 means a 10% price cut lifts usage 12-18%, so inference deflation will not rescue a debate loop that costs 12-20x tokens and loses on accuracy.
Version the harness like a dependency and benchmark it like a service, because half your agent's score lives in code you should already be planning to delete.