Engineering & Technical

The Engineer

The Signal

Kubernetes 1.35.0 runs liveness probes before your startup probe finishes.

Reproduced on k3s, minikube and kind, which puts it in upstream kubelet rather than any one distro's patches. Anything with a slow warm-up burns failureThreshold 3 during init and cycles into CrashLoopBackOff toward the five-minute ceiling. The first symptom reads as a flaky pod, not a version bug, so the debugging hours go into your application code before anyone thinks to check the kubelet.

In Play

  1. Kubernetes 1.35.0 And Go 1.27 Ship Behavior Changes You Never Opted Into

    Kubernetes v1.35.0 evaluates liveness probes before the startup probe has succeeded, and Chris Short reproduced it on k3s, minikube and kind. Pods with a long warm-up restart forever instead of failing fast: the default failureThreshold is 3, and CrashLoopBackOff widens toward a five-minute ceiling. Go 1.27, released 19 August, rebuilds the original encoding/json on top of encoding/json/v2, so serialization behavior can drift in code that never imported v2.

    Ask Clarity
    Try
  2. Scaffolding Now Moves Scores More Than Weights Do

    Nvidia's AVO wrapper took Claude Opus 5 from 30% to a perfect 183/183 on the public ARC-AGI-3 set with identical weights, per Techpresso, clearing it in 6,624 actions — about 12% fewer than the rival VISTA wrapper. Latent.Space reports a parallel result: one model scored anywhere from 52.4 to 76.2 across 106 identical tasks depending only on harness configuration. Roughly half your agent's measured capability lives in code you wrote and probably never versioned or benchmarked.

    Ask Clarity
    Try
  3. Four Agents Debating Land At Or Below Chance

    Anthropic ran the hidden-profile test — evidence split so that the facts every agent shares point at the wrong answer — on four-agent groups. Most model families decided correctly in only 17-36% of runs, per Exponential View, while a single agent handed the entire evidence base was right nearly every time. Blending several models' answers preserved only about a quarter of the good ideas one model produced, which makes a consensus stage a filter rather than redundancy.

    Ask Clarity
    Try
  4. 2027 Rack Prices And Memory Costs Reprice Your Efficiency Backlog

    Server OEMs are telling customers that GB300 and VR200 systems with 2027 delivery dates will cost about 17% more, The Information reports — near $600-700K per rack and at least $5B per gigawatt of capacity. Amazon separately moved the Echo Dot from $50 to $80 and named memory and storage component costs. Every throughput point you win through caching, batching or quantization is now worth about 17% more in capacity avoided.

    Ask Clarity
    Try
  5. Regulators Are Pricing The Missing Human-Review Step

    The Dutch DPA fined Uber €825M ($966M) on 17 August for suspending driver accounts automatically with no warning and no human review, per Techpresso — second only to Meta's €1.2B GDPR penalty. Uber's own defense, that just 126 European drivers were deactivated on ratings in 2021, shows low volume bought nothing: the regulator priced the absence of the review path. Morning Brew reports TikTok separately agreed to pay $400M to settle DOJ child-privacy claims.

    Ask Clarity
    Try

Deep Dives

Kubernetes 1.35.0 And Go 1.27: Same Interface, Different Behavior

Both bumps compile clean and pass every API check while changing how your pods restart and how your bytes serialize, so the first symptom is flakiness rather than an error.

Why it reads as flakiness, not as a version bug

The timing arithmetic is what turns this from an afternoon into a sprint. The default failureThreshold is 3, the termination grace period is 30 seconds, and CrashLoopBackOff starts at 10 seconds and doubles toward a five-minute ceiling. Each of those defaults is defensible on its own, which is exactly why nobody ever revisits them. Stacked, they stretch the feedback loop far enough that the symptom stops presenting as a deterministic fault and starts presenting as an unlucky node. The backoff is doing precisely what it was designed to do. What it was designed to do is absorb transient faults quietly, and it does not distinguish between a fault that will clear and one that will not. That is good engineering applied to the wrong failure mode, and it is the reason the first hour of triage gets spent on the infrastructure instead of the version. The numbers are the tell. Read them before assigning blame.

What to do

  1. Block Kubernetes v1.35.0 in cluster upgrade automation, including dev and CI clusters, and set initialDelaySeconds above p99 boot time on every liveness probe that sits behind a startupProbe.

  2. Add golden-file serialization tests for your 20 highest-traffic JSON types this sprint, generated on Go 1.26 and diffed against 1.27 in CI, gating the toolchain bump on that diff.

  3. Run the GA goroutineleak profile against your longest-uptime service and benchmark one allocation-heavy service on 1.27 this sprint, before the upgrade decision goes to review.

Fan Out For Capacity, Never For Judgment

Two agent results point opposite ways on scaffolding value, and the line between them decides which parts of your stack to fund and which to write for deletion.

The two different jobs people call "the harness"

Scaffolding does two structurally different things, and these results only look contradictory until they are separated. Adding steps whose output a verifier can check compounds capability: search and backtracking policy, retained reasoning, compaction, verify-before-commit. Adding voices that a reducer must reconcile averages capability away. The first is engineering. The second is a vote.

The reliability arithmetic explains why the first one works at all. A loop adds no capability. It amplifies whatever per-step reliability already exists, and below a threshold it amplifies errors instead.

Per-step reliability10 steps20 steps50 steps
0.95~60%~36%~8%
0.98~82%~67%~36%
0.99~90%~82%~61%
0.995~95%~90%~78%

Moving from 0.95 to 0.99 per step is a 2.3x improvement in 20-step success. That is where the engineering hours belong. Verification is cheaper than reliability: a passing test suite, a type check, or a fresh-context retry raises effective per-step reliability faster than any amount of prompt tuning. Derive a maximum unattended horizon from the measured number and force a checkpoint there. Same discipline as a retry budget on a downstream call.

Why the redundancy is not redundant

Given one coding task, 18 of 30 agents chose an identical git branch name, per Exponential View. That is 60% agreement on free text with effectively unbounded entropy, which measures mode collapse rather than coincidence. Every reliability pattern borrowed from distributed systems assumes independent failure domains: N-of-M voting, retry with a different seed, LLM-as-judge, ensemble guardrails. Same-family replicas do not supply that. Correlated errors do not fail noisily. They return unanimous, confident, wrong answers, the hardest incident class to detect. Move critical judging to a different vendor lineage than the generator, and track cross-vendor disagreement rate as a live signal. It doubles as a drift detector.

The refactor that keeps the legitimate reason to fan out

Convert multi-agent from a voting system into a retrieval system. Agents return structured claims carrying evidence IDs plus a mandatory "facts only I hold" field. A deterministic reducer unions the evidence with no model judgment at reduce time. One agent then decides over the full union. Instrument evidence coverage rate, the fraction of unique facts that actually reached the decision step, because that metric is the production canary for the hidden-profile failure. Prevent anchoring mechanically with sealed-bid submission. Instructing an agent to consider dissenting views does not repair a low-variance prior. Withholding the consensus does.


Which components to write for deletion

ComponentAbsorption statusHow to architect it now
Tool callingAbsorbed via in-environment RLUse native; delete custom prompt scaffolding
Context compactionAbsorbed natively in newest coding modelsAdapter behind a capability flag
Retained reasoningAbsorbed post-o1Configure, do not reimplement
Tool selection, memory, orchestrationNext in lineThin, swappable, no business logic coupling
Permissions, identity, auditAbsorption-proof: human-facingDeterministic, external, versioned

Anthropic deleted 80% of Claude Code's system prompt without losing capability, per Latent.Space. Good craftsmanship, and it makes "how much harness can we delete at constant score" a real progress metric, provided the deletions land in tranches and a separate adversarial suite stays in place, because partially absorbed guardrails fail on the tail an aggregate benchmark never samples.

What not to cite in a design review

The 106-task spread arrives with no named owner and no published methodology. Nvidia's perfect run was on the public set, with a harness retargeted to that interface and no token cost or wall-clock published. Directionally believable, quantitatively uncitable. Build a frozen suite in-house and cite that. Two measurement rules make it valid: freeze the harness to compare models, freeze the model to compare harnesses, and report actions-per-solved-task and tokens-per-solved-task next to pass rate. Measured token price elasticity of 1.2-1.8 means a 10% price cut lifts usage 12-18%, so inference deflation will not rescue a debate loop that costs 12-20x tokens and loses on accuracy.

Version the harness like a dependency and benchmark it like a service, because half your agent's score lives in code you should already be planning to delete.

What to do

  1. Freeze 50-100 real tasks from your own repos into a harness suite this sprint, then run identical weights across your harness variants and gate harness merges on the score delta.

  2. A/B every deliberation, vote, or blend node in production against a single agent receiving the union of all evidence this sprint, tracking accuracy, tokens, and p99 latency together.

  3. Move permission enforcement out of prompts into a deterministic external layer this quarter: sandboxed exec, filesystem path scoping, egress allowlist, short-TTL scoped credentials, append-only audit log.

Eleven Months Of Depreciation Erases The 2027 Rack Increase

The hike is a planning input; the useful-life assumption sitting underneath your cost-per-token model is the term that actually decides the answer.

The dominant term is not the sticker

Annualize a derived ~$3.9M rack and the ranking inverts. Five years of useful life carries about $780K a year; post-increase at the same life, about $912K. Six years lands near $760K, below the pre-increase five-year baseline. Four years is about $1.14M, a 46% increase. Stretching useful life by about eleven months fully absorbs the price move. Nvidia has been sending mixed messages on chip longevity, per the same reporting, which makes that ambiguity a bigger 2027 line item than the price. Cost-per-token quoted without a stated depreciation assumption is a guess wearing a spreadsheet.

Where the numbers come from, and how far to trust them

System-level pricing, not per-die pricing. A GB300 or VR200 SKU is a rack: GPUs, Grace-class CPUs, NVLink switch fabric, NICs, DPUs, liquid-cooling plumbing, shipped as one unit. Bakeoffs against a per-accelerator price are invalid until that vendor's rack, fabric and cooling sit in the same line. Reported figures work backward to $29-30B per gigawatt of chip-system cost; at roughly 130kW per high-density rack, that implies several thousand racks at $3.5-4M each. The per-gigawatt figure is arithmetic inference rather than a reported number, and the underlying reporting rests on two anonymous sources describing "some" flagship systems. Confirm with the OEM account team before repricing anything customer-facing.


Three levers that pay on hardware you already own

  1. Fix routing before you request replicas. A radix prefix cache is per-replica. N replicas behind a round-robin balancer push the effective hit rate toward 1/N, since the replica holding the prefix is not the one taking the request. On a dashboard that reads as "we need more GPUs." Consistent hashing on session ID or system-prompt hash fixes it at the gateway, and buys load skew, hot shards, and cache cold-start on failover.
  2. Memory-resident indexes are on the wrong cost curve. Amazon named memory and storage costs while moving the Fire TV Stick 4K Max from $60 to $85 and a 16GB Kindle from $110 to $150. Cloud memory-optimized instance families follow with a lag. Benchmark quantized embeddings and disk-backed ANN now, while it is still a curiosity project.
  3. Portability is the option premium. Port one high-volume inference workload to a non-Nvidia backend through the existing serving layer. The deliverable is the enumerated list of kernels and libraries that pin the stack to one vendor, not the benchmark.

The constraint that is not silicon

Texas halted construction on 1,800 data centers over power and water; Pennsylvania restricted expansion days later. Opposition went from 42% to 75% in twelve months, with margins of -43, -65 and -75 points among Republicans, independents and Democrats. Cross-partisan means no state's politics offers cover. If capacity lands where permits exist rather than where users are, a 35ms round-trip delta across a six-hop agentic chain is over 200ms of added p99. Parallelize tool calls, cache prompts hard, colocate the vector store with the model endpoint rather than the app tier. Log kWh and dollars per million tokens per model and per endpoint; power and water are the stated grounds for every moratorium, and procurement questionnaires are coming for numbers almost no team collects.

Silence on chip longevity is a bigger line in your 2027 cost model than the price increase everyone is reacting to.

What to do

  1. Classify every 2027 GPU capacity commitment as locked-price or subject to pass-through this quarter, then re-run your dollars-per-million-tokens model at 4-, 5-, and 6-year useful life and publish the sensitivity band instead of a point estimate.

  2. Add prefix-affinity routing at the model gateway before requesting more replicas, and measure the cache hit-rate delta on replayed production traffic rather than synthetic uniform prompts.

  3. Benchmark quantized embeddings and disk-backed ANN against your fully in-memory index before the next memory-optimized reserved-capacity renewal.

The bottom line

These items share one shape: the name you depend on stayed the same while the thing behind it moved. A minor version, an import path, a weights reference, a scaffold config — none of them predict behavior anymore, and none of them fail loudly when they drift. What breaks is the reflex that a green build plus an unchanged dependency list proves nothing changed underneath you. Pick the three layers you still resolve by name rather than by pinned digest or measured score, and put a golden-file or frozen-task gate in front of each one.