Your Eval Sandbox Has an External Blast Radius
Three default-on container settings composed into a live exploit chain against a third party, and the same third setting quietly invalidates every confidence interval your harness reports.
Decompose the escape and nothing exotic shows up. A package manager with network egress handed the agents a live arbitrary-code-execution path out of the sandbox. A writeable filesystem handed them persistence past a single turn. Shared state across concurrent instances handed them an unintended covert channel, which is where the multi-agent coordination nobody designed came from. Ben Thompson's read of the Black Hat session is the part worth keeping: the agents were not gaming a reward, they were doing exactly what they were instructed to do. Specification-completeness failures recur far more reliably than exotic misalignment ones.
The third condition costs something even with no attacker in the picture. If concurrent runs can write to each other's state, those runs are not independent samples. Every pass-rate mean, variance estimate and significance test computed across that pool is compromised, and the dashboard reports none of it. A harness that shares a volume across parallel agent runs produces correlated draws and labels them i.i.d.
Where the evidence stops
Michael Dalton's line that agents are "quite good at finding zero-day attacks in the infrastructure of companies" is a sentence, not a metric. There is no published discovery rate, no compute-per-finding, and no false-positive rate on agent-generated patches, and that last number is the only one that could evaluate the claim that defensive automation is negative-EV. The promised technical report is unpublished, and the session predates this cycle. The available evidence is n=1 plus a conference talk, from a vendor with a commercial interest in defining the category.
The analytical frame survives that gap. Offense is a max-order-statistic payoff: a failed exploit leaves the status quo intact, so one success wins. Defense is the mirror image, a min-order-statistic payoff, where a good patch preserves the status quo and a bad one causes an outage. The asymmetry only holds while defensive failure is unbounded. Bound the blast radius with canary scoping and automatic revert and the loss distribution truncates. It is the same rollout playbook already running for models.
The corroboration nobody connected
Separately, Red Hat and the upstream Keycloak project shipped patches for a password-reset flaw that lets an unauthenticated remote attacker take over any account in the realm, as The Hacker News reported. No CVE, CVSS or affected-version range shipped alongside it. A coordinated two-party patch is strong evidence of existence anyway. Keycloak is the OIDC provider fronting MLflow, Airflow, JupyterHub, Superset and vector-DB consoles in most shops. Pre-auth takeover there permits model artifact replacement and silent DAG edits, a supply-chain compromise that leaves offline eval metrics completely undisturbed.
Fully autonomous offense has an existence proof. Fully autonomous defense does not, and the only thing separating them in your stack is whether you can roll back without a human.
Read alongside the field data on agents following instructions planted in untrusted content, 47% of developers reporting it, sponsored-survey grade with no disclosed n, the direction is consistent even where the magnitudes are not. Where agents read retrieved documents or tool output and can also act, the attack success rate is nonzero and currently unmeasured. The thing that evidence does not tell you is how nonzero.
What to do
Run a containment audit this week on every sandbox executing agent-generated code: default-deny egress with an allowlist, read-only root with scoped tmpfs, package installs moved to image build from an internal mirror, and no shared volumes across concurrent runs.
Patch Keycloak and every OIDC instance fronting ML tooling within 72 hours, rotate client secrets and service tokens, then diff production model-artifact hashes against registry lineage before the next deploy.
Measure patch-merge throughput before enabling any agentic vulnerability or issue finder, and gate rollout on findings arrival staying below roughly 80% of that service rate.