Your Noise Floor Lives in the Sharding Config
Three headline results were reported with no model of the variance in the system that produced them, and each one has a cheap test that settles the question.
Why a sharding change is a treatment, not a config
Floating-point addition is not associative. The order in which partial sums get reduced changes the result in the low bits. Move a run from 8-way to 16-way tensor parallelism for throughput and that reduction order changes at every layer on every step, and the divergence compounds through optimizer state. AI News relays the finding that this numerics-level nondeterminism produces run-to-run variance nearly as large as changing the initialization seed or the data order.
Teams change topology more casually than any other variable, usually for cluster availability or a throughput win on a new node shape. Any ablation whose arms span that change is comparing method plus device layout. Evidence caveat: this arrives as a relayed summary rather than a paper with ablation tables. Treat it as a hypothesis expensive enough to test, not a settled result. The test is three to five repeats of a configuration already in rotation.
The protein headline is a max statistic
Pivot 5's decomposition of Anthropic's August 19 technical reports is the clearest public example of a unit-of-analysis error. Of 1,320 tested designs, 354 bound. That is 26.8%, or roughly 31% excluding two dead targets. At about 88 designs per target and a ~27% per-attempt success probability, the probability of at least one binder per target is effectively 1.0. The 14-of-15 target-level figure is therefore arithmetically implied by the design-level rate rather than independent evidence of it. There was one replicate per model-target combination, no human-expert control arm, and the endpoint was binding only, not affinity, specificity or stability. The plannable number is 27%.
The gate that failed was the one guarding the money
The detail that transfers furthest past the biology: folding-model confidence scores flagged neither of the two complete target failures. Two targets consumed synthesis budget with zero in-silico warning. Strip the biology and this is the most common silent failure in production ML, an uncertainty signal with respectable average AUC that collapses in the tail and then gets wired into an automated gate because the average looked fine. Average calibration is not the property the gate needs; tail calibration is, and it is almost never what gets measured.
| Claim | Reported number | Missing control | Cheap test |
|---|---|---|---|
| Numerics drives run variance | Rivals seed and data order | No ablation table, no CIs | 3-5 repeats varying only seed and shard layout |
| Designed binders | 14 of 15 targets | n=1 per combo, no expert baseline | Report per-attempt rate alongside best-of-N |
| Confidence as a triage filter | Used pre-synthesis | Tail behavior never measured | Reliability diagram, ECE, worst-decile coverage |
The confound that is physical
TLDR Hardware supplies the third instance, and it is the one nobody logs. Liquid cooling moved the thermal control loop outside the server: cold plate, manifold, CDU, facility heat rejection. Pumps and valves are explicitly designed to compensate for developing faults invisibly, so component temperature confirms operation within limits while revealing nothing about remaining margin. Quietly throttled hosts make latency non-stationary across nodes, which violates exchangeability in randomized latency tests and produces benchmark runs that will not reproduce later.
If the ablation delta is smaller than the floating-point noise floor, the comparison has no information in it.
All three sources agree that variance is systematically under-instrumented relative to effect size. They differ in evidence grade. The protein numbers come with full counts anyone can recompute, the numerics claim does not, and the thermal argument is mechanism rather than measurement. The common root cause is that all three effects are produced by the infrastructure layer, which sits outside every experiment log a team keeps.
What to do
Run your champion training or fine-tuning config 3-5 times this sprint, varying only seed, data order and shard topology, and publish the resulting metric band as a required field in the ablation template.
Audit every model confidence score that triggers routing, abstention or spend this sprint with a reliability diagram, ECE and conditional coverage on the worst-outcome decile; demote any gate below AUC 0.6 on real failures to logging only.
Add device clock, power draw, throttle reason and host ID to serving telemetry before your next latency experiment, and block or randomize arm assignment by node.