Your Autoscaler Assumes Infinite CPU. That Assumption Just Died.
Spot capacity vanished, reservations get rejected for the wrong SKU, and the fastest-growing consumer of cores is the agent fleet you just moved off developer laptops.
The squeeze is upstream, and you can't engineer around it
The mechanism sits in the fab, not your cluster. Anthropic's Katelyn Lesse describes three factory bottlenecks compounding at once: TSMC lines where GPUs outbid CPUs and where fabless AMD competes for the same allocation; memory wafers where makers favor high-margin HBM over the DRAM every CPU server needs; and Intel's yield-constrained fabs. Analysts expect CPU headroom to return before memory does, and that is still multiple quarters away. The tempting escape — move to Arm or another vendor — buys SKU flexibility when a reservation is rejected for the wrong type, but hyperscaler Arm silicon largely comes off the same TSMC lines. It hedges your allocation, not the constraint.
Agents are CPU workloads with a GPU front-end
The demand side is what lands on your roadmap. An agent request does two things: it generates code (GPU-bound inference) and then runs tools — compile, test, lint — which is plain general-purpose CPU. Uber and Ramp already moved agents off laptops onto dedicated cloud instances; Uber's agentic requests rose ninefold in six months. In systems terms, every agent session is a CI job nobody scheduled. Reinforcement-learning rollouts at the labs add to it: teaching a model to search means the model actually executes software, so every rollout burns CPU. That is the mechanism behind AI data-center CPU:GPU ratios sliding from 1:8 toward 1:1 — eight times the host CPU per deployed GPU at the limit.
The cost signal reinforces this from another angle. A single four-hour agent task — generating a marketing course — consumed 22% of a weekly allowance on a $100 Codex plan. The binding constraint there isn't dollars per token; it's plan capacity per engineer. Both stories describe the same thing: agent work consumes a scarce, lumpy resource in units nobody is metering yet.
In your stack
Autoscaling inherits a new failure mode. A capacity-denied scale-out should page someone, not silently retry, because your SLOs now depend on a provider decision you don't control. Single-SKU node pools are a single point of failure; diversify instance families, sizes, and AZs, and ship multi-arch images to widen the set of acceptable SKUs. Scale-in is now a one-way door — capacity released at 3am may not return for the 9am ramp — so hold a warm floor for scarce SKUs and treat the idle cost as insurance. Load shedding and queueing become capacity tools, not just resilience tools.
Put agent tool execution on its own CPU pool with per-task quotas, isolated from production, and attack CPU-seconds per task with the unglamorous CI work: remote build caches, test-impact analysis, diff-scoped linting, pre-warmed sandbox images. Caching trades CPU for network and storage, the right trade when networking is the one primitive not in short supply. Before buying anything, reclaim what you already pay for — the gap between requested and used CPU in Kubernetes is likely the cheapest capacity you'll find this year.
Calibrate the alarm honestly. This reporting is direction-strong and quantification-thin: much of it is dinner-table anecdote and anonymous sourcing from the heaviest compute buyers, resent from an issue roughly two weeks old, with no hyperscaler confirmation. A mid-size shop on steady reserved instances may feel this as higher prices, not denied capacity. The fastest calibration is your own data — pull 90 days of spot fulfillment and capacity-error metrics before signing any multi-quarter commitment.
What to do
List every spot- and autoscale-dependent workload (CI runners, batch/ETL, agent sandboxes) and pull 90 days of spot fulfillment and capacity-denied metrics this week, before committing to any reservation.
Reserve a conservative CPU floor against a 12-month forecast and negotiate flex or break terms on anything prepaid for December-onward capacity.
Move agent tool execution onto an isolated CPU pool with per-task quotas this sprint, and cut CPU-seconds per task with remote build caches and test-impact analysis.