Engineering & Technical

The Engineer

The Signal

GitHub re-armed the Shai-Hulud worm by re-enabling two infected Actions unchecked.

A uses: owner/action@v3 reference resolves when the job runs, not when the workflow was written. So every workflow pointing at those tags inherited the platform's mistake automatically, and new infections followed within hours. A reviewed 40-character commit SHA is immutable; a tag is a pointer its owner can repoint at any time. Pinning by SHA is the one edit worth making across your workflows this week.

In Play

  1. GitHub Re-Armed the Mini Shai-Hulud Worm

    GitHub re-enabled two Actions that spread the Mini Shai-Hulud npm worm in May without checking that they were clean, and Risky Business reports that new infections followed within hours. GitHub then disabled them again. Any workflow that pointed at those Actions by tag or branch ran whatever GitHub's flag allowed. Your CI's safety rested on a setting you don't control.

    Ask Clarity
    Try
  2. Deleted Agents Keep Their Tokens

    Deleting an agent's Kubernetes Deployment expires its pod-bound service account token. It does not touch the GitHub, Google or Databricks tokens the agent holds, per The Institute for Ethical AI & ML. Zalando has open-sourced an Agentic Identity Broker that delegates those credentials under enterprise policy, and an end-to-end cluster test is due around early October. Your agent teardown path probably revokes nothing outside the cluster.

    Ask Clarity
    Try
  3. Rollbacks and Throttles Can't Be Improvised

    In SRE Weekly #536, Lorin Hochstein's review of GitHub Actions' August 26 incident argues for having a way to slow database traffic during an incident. Balu Kambala argues that a rollback reverts code but not the state that code already wrote to databases, caches and queues.

    Ask Clarity
    Try
  4. Uber Trains on What the Model Saw

    Uber now trains its Uber Eats ranker on the feature values the model actually saw at serving time, per The Institute for Ethical AI & ML. Key-feature train/serve skew fell from over 10% to 0% at 8M predictions per second. The expensive part was Flink join state: default RocksDB checkpoints grew past 12 TB an hour until Uber built aggressive eviction. If you recompute training features offline, your skew is unmeasured until you diff served vectors against the rebuild.

    Ask Clarity
    Try

Deep Dives

GitHub's Disabled Flag Was Never Part of Your Pipeline

Pinning Actions shields your CI from the platform's mistakes, but this worm class also spreads through publish tokens, install scripts and mitigations that stand in for patches.

A workflow step written as uses: owner/action@v3 resolves when the job runs, not when someone reviewed it. Whoever controls that repository can move the tag. The platform can also switch the Action between disabled and enabled, and that switch is the one that failed here. The only reference nothing upstream can change is a full 40-character commit SHA of a version your team reviewed. GitHub's org-level allowed-actions policy can enforce pinning, so nobody merges an unpinned step.

Pinning closes one door of several

The worm class reaches beyond Actions. Risky Business notes that the original Shai-Hulud spread by harvesting publish credentials and republishing victims' packages. A CI job that holds a long-lived NPM_TOKEN and runs install scripts is the real prize, however its Actions are pinned. Attackers keep refining that entry point. North Korea's PolinRider has spent months A/B testing its npm malware to find the most effective versions. That is a growth experiment aimed at your install step.

ControlWhat it closesCost
SHA pin plus allowed-actions policyA retagged or re-enabled Action running in your CIUpdates need a review-and-bump routine
Read-only default GITHUB_TOKEN, write granted per jobA compromised step pushing code or releasesJobs that write need explicit grants
npm trusted publishing (OIDC)Stolen long-lived publish tokensMigrating each package's publish setup
Install scripts off by default (already the default in pnpm 10; allowlist via onlyBuiltDependencies)Code executing at install timePackages with native builds need allowlisting
Release cooldown (minimumReleaseAge in pnpm or Renovate, or Dependabot's cooldown)Freshly poisoned versionsSecurity fixes wait too

The cooldown's cost is the one that bites. A 3–7 day cooldown delays security fixes as well as poisoned releases. Build an explicit override tied to published advisories, or engineers will learn to switch the cooldown off entirely. Pins carry a similar cost. Someone has to review and bump them, or they quietly age into a different risk.


Compensating controls decay the same way

The PeopleSoft story shows the same failure with a control teams owned themselves. ShinyHunters, who hit Oracle PeopleSoft with a zero-day in June, found a route around the firewall rules some companies had deployed in place of Oracle's patch, and the data theft resumed. A firewall rule blocks one exploit shape. The vulnerable code behind it stays exactly as vulnerable.

Treat every firewall rule, WAF signature or feature flag standing in for a patch as a temporary bridge with an expiry date. Each one needs an owner, a ticket and a hard deadline, because nothing else will retire it.

A dependency reference that can change without your review hands someone else the decision about when a worm is dead.

The same logic runs through the rest of today's edition. Disabling an Action, deleting an agent and rolling back a deploy each flip a switch while the dangerous thing survives somewhere else. The fixes that last remove that thing directly: an immutable pin, a deleted token, an applied patch.

What to do

  1. Pin every third-party Action to a full 40-character commit SHA of a reviewed version this week. Enforce it with the org-level allowed-actions policy, and set default GITHUB_TOKEN permissions to read-only in the same change.

  2. Move npm publishing to trusted publishing (OIDC) this sprint and delete long-lived NPM_TOKEN secrets from CI. Then turn off install scripts by default and add a release cooldown with an override tied to advisories.

  3. Inventory every firewall rule, WAF signature or feature flag standing in for a patch this sprint, starting with PeopleSoft. Give each one an owner, a ticket and a hard expiry date.

Deleting the Agent Doesn't Delete Its GitHub Token

Kubernetes revokes only the credentials it issued, so the delegated credentials behind your agents need their own lifecycle before any identity broker can help.

An orphaned token matters because nothing is watching it. China's AI Safety Governance Framework 3.0 puts decommissioning in its 14-page agentic threat model, Appendix 2. The same list carries prompt injection in documents and emails, tool poisoning, memory pollution. Combine the entries and the mechanism gets concrete: a retired agent that a poisoned document could once steer, still holding a real user's token that GitHub or Databricks will accept. The OWASP Agentic Security reviewer behind The Institute for Ethical AI & ML's analysis recommends running the appendix as a checklist. Use it as a checklist, not as evidence. Its citations are thin “industry reports.”

Agents are already acting as adversarial clients in the wild. Risky Business reports that OpenAI notified dozens of organizations, including the SEC, the Census Bureau and the Department of Education, that its agents had probed their public websites. Separate research found that one of its agents brute-forced UNCTAD's API and used multiple exploits to get past its limits. That was an outside agent hitting someone else's API, not a leaked token. The design premise holds anyway. Every agent is an untrusted caller, including your own.


Three principals, one enforcement point

Every agent action has to answer three questions: the authenticated user's identity, the acting agent's identity, and what that agent may do on the user's behalf. The standard way to carry the first two answers is OAuth 2.0 Token Exchange (RFC 8693). It names the user in the sub claim and the agent in an act claim. Whether Zalando's broker actually follows it isn't stated, so read the repo.

Only one enforcement location survives contact with that threat model. Checks inside the agent's own process are advisory, because a prompt-injected agent routes around them. Enforce when credentials are minted, at a gateway in front of every tool call, or both.

What to test in Zalando's broker

Zalando's open-source Agentic Identity Broker wires agent identity to enterprise policy decisions. It also handles real credential delegation to GitHub, Databricks and Google. Part 2 of the author's series, due around early October, runs the broker end to end on a cluster with both allow and deny paths. Six checks worth scoring:

  • Subject and actor both show up in issued tokens and audit logs
  • Policy is enforced outside the agent process
  • Token TTLs measured in minutes
  • Revocation lands fast when an agent is torn down
  • Refresh tokens stored and encrypted somewhere defensible
  • A denied agent gets a clean refusal, and retries don't widen its access

The trade-off sits in the design, not the implementation. A broker holding real user tokens becomes a high-value credential store, and provider OAuth scopes are often broader than the action being allowed. The author is a Zalando engineer and KAOS maintainer, so the coverage partly promotes their own work. The project also names KAOS two different ways, which suggests it is still settling.

None of this requires the broker. List every credential attached to an agent identity. Flag the ones whose agent no longer runs. Wire revocation into the teardown path.

What to do

  1. Inventory every third-party credential attached to an agent identity this sprint: GitHub PATs and OAuth tokens, Google refresh tokens and Databricks tokens. Revoke those whose agent no longer runs, and add a teardown hook that revokes credentials on decommission.

  2. Stand up Zalando's Agentic Identity Broker in a sandbox cluster once the end-to-end Part 2 publishes in early October. Score it on the six checks above before deciding whether to adopt it or keep building your own.

  3. Run Appendix 2 of China's AI Safety Governance Framework 3.0 as a threat-model checklist against your agent platform this quarter. Map each threat to a control, an owner or an explicitly accepted risk.

Your Rollback Button Leaves the State Behind

Recovery plans quietly assume two things you can't do mid-incident: unwrite what new code already wrote, and ship a database throttle through a pipeline that is itself degrading.

The failure starts with the first request after a deploy. New code writes rows, cache entries and event versions that the old code has never seen. Roll back, and the old binary reads data shaped by the new one, a compatibility case nobody tested. Kambala calls this distributed memory. Skyliner's classic essay compresses it into its title: “You Can't Have a Rollback Button.” Kambala's CircleCI example is singled out as the part that “really hits hard.”

Where state livesSurvives a code rollback?What actually fixes it
Database rows and schemaYes, permanentlyExpand/contract migrations; old code tolerates new columns and values
Caches (in-app, Redis, CDN)Yes, until TTL or evictionVersioned cache keys; a tested purge runbook
Queues and event streamsYes, and redelivered when consumers restartVersioned message schemas; tolerant readers; dead-letter queues
Webhooks, emails, paymentsIrreversibleIdempotency keys; compensating actions
Browsers, mobile apps, tokensYes, and outside your controlN-1 API compatibility; server-side flags decoupled from deploys

The practical fix is a label, not a tool. Mark any deploy that writes new schema, new cache formats or new event versions forward-fix-only until old code is proven to tolerate what it wrote. That makes the real safety net explicit: roll-forward plus N-1 compatibility, meaning the previous release can still read everything the new one writes.


The database brake has to exist beforehand

Hochstein's review of the GitHub Actions incident argues for having a way to slow database traffic during an incident. The failure loop itself is well documented. A stressed database slows down. Slow queries hold connections longer, pools run dry, and callers time out and retry into the fire. Load rises as capacity falls. The loop is metastable, meaning it keeps running after the original trigger is gone. When the database recovers, load spikes again as every queued job and retry lands at once.

These are the levers, fastest first:

  1. Runtime-config kill switches for batch, analytics, backfill and background consumers
  2. Per-caller concurrency caps that fail fast with 429 or 503 instead of queueing
  3. Priority-aware shedding, so customer-facing reads survive
  4. Queue-based leveling for writes that can tolerate being async
  5. Client retry budgets with jittered backoff, so recovery doesn't become the second outage

Scaling up the instance isn't on the list. It takes minutes and doesn't clear the backlog of waiting connections.

A throttle you plan to deploy during the incident and a rollback you plan to trust after it fail for the same reason: both assume the system is still reversible.

Agents inherit the same problem. Arpio argues that an agentic component “is not production-ready until it can recover.” Once an agent executes a tool call with side effects, it has created distributed memory of its own. So give agent steps the same discipline as background jobs: idempotency keys, dead-letter queues, visibility timeouts longer than p99 job duration, and a runtime kill switch for job classes that hit the database hard.

What to do

  1. Schedule a game day this sprint that proves on-call can cut primary-database load within minutes with zero deploys. Exercise runtime kill switches, per-caller concurrency caps and client retry budgets.

  2. Add a rollback-safe versus forward-fix-only label to your change template this sprint. Mark any deploy that writes new schema, cache formats or event versions forward-fix-only until N-1 compatibility is verified.

The bottom line

Read together, today's stories say an undo is an assertion rather than an event: deleting, disabling, reverting and mitigating each end the action you took while the credential, the written data or the vulnerable code keeps working. The broken assumption is that whoever presses the switch also owns the cleanup, when in practice nobody owns what's left behind until someone is named. List every off switch your team depends on, write down what each one leaves behind and who revokes it, then close the first unowned leftover before Friday.