The Wrapper Is the Treatment and the Model Is the Control
Two independent results put a bigger score swing in the orchestration layer than in the weights, and the scaffolding that produces it is being absorbed upstream while you build it.
Both designs are paired, and neither used the statistics that buys
Granularity check first, because it decides what you can compute on Harness-Bench. Neither 52.4 nor 76.2 is an integer multiple of 1/106, where one task is worth 0.943 percentage points, so the scoring is either partial-credit or averaged over several runs. Latent.Space never says which. That single omission decides whether a per-task paired statistic is available at all.
At n=106 with accuracy near 0.5, an unpaired 95% binomial interval has a half-width of about +/-9.5 points. The 23.8-point extremes clear that comfortably. Two mid-pack harnesses separated by less than roughly 13 points are indistinguishable under a two-proportion z-test. The design is paired, though: same model, same 106 tasks. McNemar's exact test on discordant pairs recovers far more power from the identical data. Benchmark harnesses internally with independent-proportion tests and most of the sample goes in the bin.
What the 100% does and does not establish
Techpresso reports AVO clearing all 183 levels in 6,624 actions, about 12% fewer than the rival VISTA wrapper. Three caveats travel with it. The score is on the public set of a benchmark whose entire design premise is resistance to memorization, and a wrapper author can see those games, so what it establishes is that scaffolding can saturate an eval. That is a weaker claim than solving novel reasoning. The 12% action-efficiency edge arrives with no per-game distribution and no seeds, so adopt the metric and discount the ranking. The finding that survives contact with production: AVO was originally built to tune GPU kernels and was retargeted by swapping its tool interface. The reusable asset is the control loop, meaning planning, retry, state tracking, budget management, not the domain tooling.
The one piece of clean math to plan against
Latent.Space's compounding-error model is analytic, which makes it the most reliable number in either report: 0.95 per-step reliability over 20 steps yields 35.9% task success. Inverted, it reads as a spec sheet. Eighty percent success at 20 steps requires 98.9% per step; at 50 steps it requires 99.6%. Moving per-step reliability from 95% to 99% takes 20-step success from 36% to 82%.
That is where the engineering budget belongs, and it explains the other reported lift: retained reasoning plus compaction tripled GPT-5.6 Sol on ARC-AGI-3 from 13.3% to 38.3%. Near n=60 that is roughly 8 to 23 solved tasks, and it clears p<0.001 under McNemar with room to spare. The two interventions were reported as a bundle with no ablation and no token cost, so the effect is real but you cannot tell which half did the work or whether it is affordable. Mechanistically, neither adds capability. Both prevent state loss that otherwise surfaces as step failures.
Scaffolding is a depreciating asset, so instrument it that way
Compaction has already crossed from harness into weights. GPT-5.1-Codex-Max is described as the first model natively trained to operate across multiple context windows through compaction, and Anthropic deleted 80% of Claude Code's system prompt and framed it as progress. Named next in line: multi-agent orchestration, tool selection, memory. Feature-flag every scaffold component so it stays ablatable, run a deletion sweep on each model upgrade, and keep whatever stays deleted at constant eval. Gate those deletions on edge-case canaries rather than aggregate score, because partial absorption produces tail regressions a difficulty-averaged benchmark will never surface.
Vendors ship model cards. Nobody ships a harness card - so co-evolution gains are documented nowhere, and a pinned harness is the only instrument you control.
What to do
Log harness_hash beside model_hash on every eval row this sprint - tool manifest, compaction policy, retry semantics, permission config - and re-run your most recent model comparison with the harness pinned before the next selection decision.
Run a 2x2 ablation of reasoning-trace retention against context compaction on your own agent benchmark within two weeks, reporting accuracy and tokens-per-solved-task for each cell.
Add a scaffold-deletion sweep to the model-upgrade checklist this quarter, toggling each prompt section, router heuristic and memory layer off, with promotion gated on edge-case canary pass rather than aggregate score.