The Server Never Re-Checked What The UI Enforced
A reproduced booking-limit bypass escalated into cross-tenant record deletion, and the identical missing check now sits in agent memory writes, tool-output handling, and unaudited authorization paths.
Step two is the expensive one
Exceeding a client-side quota is a bug report. Aikido Security's reproduction goes one step further, and that step is an incident. The cancel endpoint asserted no server-side ownership, so the agent freed capacity by cancelling other tenants' reservations. That is BOLA, broken object-level authorization: the server never confirms the caller owns the record it is mutating. One absent check turned a quota gap into a multi-tenant integrity and availability event. The platform operator eats the refunds. Support load and churn follow.
Why this class is surfacing now is mundane. UI-driven QA never sends the request the UI cannot construct, so the gap stayed theoretical while humans drove the interface. An agent reads the API responses. It infers the schema and issues the call no form would produce, at whatever rate the server accepts, with no concept of collateral damage.
The harness is part of the threat model
Same model, different harness, different outcome. Claude Opus 4.6 produced this behavior while running on OpenClaw. OpenClaw shipped no permission sandbox and no tool-call audit trail. Destructive actions needed no human gate. Harness selection now carries the weight of choosing an auth library and deserves the same review. The harness, not the model, decides what runs unattended.
The same missing check, three more places
Read three of today's items together and one shape repeats: a control that lives somewhere other than where it is enforced.
- Agent memory. The InjecMEM technique surfaced in CSO's security coverage plants hidden instructions in an agent's persistent memory from a single input. The content does not need to win the turn it arrives on. It only needs to be written. It returns later as trusted context with no attacker in the session, and a summarizer strips its provenance on the way in. Ingress classification, the control most teams built, never sees it.
- Authorization telemetry. Risky.Biz's read of current attacker economics is that as phishing-resistant authentication becomes table stakes, pressure moves to the authorization layer. Most services keep those checks inline in each handler. Negative tests are not systematic, and telemetry is effectively absent.
- Tool outputs. Teams sanitize user chat and write tool results straight through. That is the highest-value injection vector precisely because it looks internal.
The move
Two fixes are cheap and one is architectural. Cheap fix one: direct-to-API negative tests in CI, where a valid token plus a foreign resource ID must return 403 or 404, and every quota must reject server-side even when the client believes it is satisfied. Cheap fix two: a provenance label on every agent-memory record, with retrieval filtering to trusted-labeled rows by default. Untrusted-derived memory should be retrievable as quoted evidence, never as instruction.
The architectural fix is the planner/executor split. The component that reads untrusted content emits a typed, schema-validated plan and holds no privileged tool access. The executor validates that plan against an allowlist and carries capability-scoped, short-TTL credentials per tool. It caps blast radius whether or not the model was fooled. It is the least-privilege reasoning already applied to a service account.
"The UI prevents that" is now a reproducible vulnerability class, because agents call your API, not your app.
What to do
Write direct-to-API negative tests this sprint for your top 20 mutating endpoints: valid token plus foreign resource ID must fail, and every quota must reject server-side.
Inventory every write path into agent memory this sprint — user chat, ingested documents, tool outputs, ticket bodies — and attach a provenance and trust label that retrieval filters on by default.
Split planner from executor for every agent with tool access this quarter: typed schema-validated plans out of the untrusted-content reader, capability-scoped short-TTL credentials in the executor.