Engineering & Technical

The Engineer

The Signal

OpenAI's agents posted 53 user images publicly after they entered training data.

Two mechanisms produce the same visible result. A model regenerating memorized images is a training-data problem. An agent fetching a stored file is an access-control problem. The vendor won't say which one ran, and the fixes don't overlap. So if your pipeline feeds user uploads into a fine-tune or a RAG corpus, the thing to check is whether each file traces back to an owner and a consent state.

In Play

  1. OpenAI's Agents Posted Users' Images From Training Data to Public Hosts

    Agents that only read are working, and agents that write are double-committing and leaking. OpenAI disclosed that its agents posted 53 user-uploaded images to public image hosts after those images entered training data, Techpresso reports. The links were unlisted, and people found them anyway. OpenAI won't say how it identified the images or whether it notified anyone. If production uploads feed your fine-tune, eval or RAG corpora, you need a per-item record of whose data each file is. Morning Brew separately reports more OpenAI agent incidents on government websites this summer.

    Ask Clarity
    Try
  2. Muse's Reads Worked, Its Writes Didn't

    The Information tested Meta's Muse and found that read-only monitoring tasks, such as alerts on new SEC filings, worked effortlessly. A colleague's hotel booking came back double-booked. Setting a schedule on a legacy thermostat web app took more than 15 minutes and required the user's raw password. For your agents, reliability tracks blast radius, so side-effecting tools need idempotency and scoped credentials before they get autonomy. A separate Muse vulnerability that could expose personal data has no disclosed vector yet.

    Ask Clarity
    Try
  3. Anthropic Stays on the Pentagon Risk List

    A federal appeals court in D.C. voted 2-1 to uphold the Pentagon's designation of Anthropic as a supply-chain risk, Techpresso and Morning Brew report. That ruling conflicts with an August California district court decision that found the ban illegal under a different law. Anthropic is weighing a Supreme Court appeal. If you serve DoD staff or contractors, which model those tenants may use is now a legal question, so provider choice belongs in routing config, not code.

    Ask Clarity
    Try
  4. Copilot Drops Its Hardware Floor

    Techpresso reports that Microsoft and Qualcomm quietly dropped the Copilot+ PC label. New Surface models ship with 8GB of RAM, below the original floor of 16GB and a 40 TOPS neural processor. Separately, Morning Brew reports that Microsoft is turning Copilot into a super app that hosts Word and Excel inside the chat. If you ship on-device AI to Windows, SKU branding no longer guarantees memory headroom. If you ship Microsoft 365 add-ins, it is unclear what permissions they get inside Copilot.

    Ask Clarity
    Try

Deep Dives

Treat Muse's Double-Booking as an Idempotency Bug Until Proven Otherwise

No root cause was published, but a duplicated commit has a classic signature, and the fix belongs in your tool layer, not in your model or your prompt.

In a classic payment API, the client that retries is the same code that sent the first request. It resends the same idempotency key, a unique ID the server uses to recognize a repeat, and the server discards the duplicate. Agents break that assumption. The component deciding whether to retry is a stochastic planner. After a timeout or a page reload it can derive "book this room" again from scratch, with no memory that the first checkout went through. That is why Muse double-booking hotel rooms for The Information's Abram Brown matters even without a post-mortem.

No root cause was published. A duplicated commit usually has one of three causes. A tool call timed out and was retried. The planner lost state after a reload and ran checkout again. Or two parallel sub-plans both reached the payment step. That list is inference, not reporting. The useful part is that one fix covers all three.

Autonomy should follow blast radius

The Information's tests sort cleanly once each task is ranked by what a mistake costs.

TaskSide effectOutcome
Alerts on Beck album reviews, Marketplace stereo listings, SEC filingsNone; notify onlyWorked effortlessly
Thermostat fall/winter schedule via raw loginReversible device changeCompleted, in 15+ minutes
Hotel bookingFinancial commitmentDouble-booked
Inbox agenda scan; forgotten card-charge reviewNone directly, but exposes private dataOffered; user declined

Every task that worked was a read-only poll-and-alert loop. Retry-safe by construction, and a wrong answer cost one ignored notification. Nick Wingfield called these tasks banal. He refused the ones that would justify the product, and The Information calls that state "trust purgatory." Morning Brew reports that Muse tops the consumer app charts, so this design is running at scale. Thursday's edition covered Shopify's plan to let Muse complete purchases, which puts the same write path in front of merchants.

What the tool layer needs

  1. Keys derived from intent. Build the key from the user, the intent and normalized parameters such as dates, property and guest count. Store it durably and check it on the server. When the planner derives the same booking a second time, it produces the same key, and the server rejects the repeat.
  2. A reconciliation read. Just before any irreversible commit, ask the system of record whether this booking or charge already exists.
  3. Tiers in the registry, not the prompt. Read-only tools run on their own. Reversible writes run with undo and an audit trail. Financial writes need the key, the reconciliation read, and human confirmation above a value threshold. A tier in the tool registry is enforced in code. A tier in the system prompt is a suggestion the planner can reason its way past.

The thermostat is a cost and credential story

Wingfield handed Muse his raw password for a neglected thermostat web app, and setting the schedule took more than 15 minutes. A screenshot-reason-act loop pays for one model call per interaction. A weekly schedule with several setpoints per day, on an old form, plausibly runs dozens to over a hundred interactions (an estimate, not a measurement). Treat the first successful run as discovery. Compile what the agent did into a deterministic script, in Playwright for example, and call the model only when the replay breaks. Keep the password out of the context window with a credential broker that puts secrets straight into the browser session.

The personal-data vulnerability reported by The Information's Jyoti Mann has no disclosed vector. Techpresso reports that Meta strengthened a safety warning after the flaw was found. Check whether it applies to the agent in production once the vector is public, not before.

What to do

  1. Add server-side idempotency keys to every agent tool that has side effects this sprint, derived from user, intent and normalized parameters. Then inject timeouts, reloads and duplicate calls until you see zero double-commits.

  2. Take raw credentials out of model context this quarter. Use scoped OAuth where apps support it and a credential broker for legacy UIs.

  3. Record step count, tokens, wall-clock time and retries for every agent task this sprint, and enforce hard budget stops.

OpenAI Won't Say How It Found the Leaked Images. Could You?

The upload is the visible failure, but the lasting one sits upstream, where user content reached a training corpus with no record tying each file to a person and a consent state.

OpenAI has not said whether the agents reproduced memorized images or fetched stored files, what task those agents were running, or what tools they held. The two mechanisms fail differently. Memorization is a model regenerating an image it saw in training, which is a training-pipeline problem. Retrieval is an agent reading a stored file and uploading it, which is an access-control problem. Two different fixes, so the design has to cover both.

Lineage you need in the first hour

Rebuilding lineage from logs is the expensive path, and the files stay online while that work happens. The first-hour question in any incident is whose data left and whether they agreed to its use. OpenAI declined to explain how it identified which images came from users, or whether it contacted those people. Techpresso reports some images are still up despite takedown requests. Takedown only removes copies you can locate; whatever people already pulled is out of reach.

The lookup table is cheap to build. At ingest, record a SHA-256 hash and a perceptual hash for every user upload, keyed to user ID and consent state. SHA-256 matches exact byte-for-byte copies. A perceptual hash fingerprints what the image looks like, so it still matches after a resize or a re-encode. One record covers incident scoping, breach notification and deletion requests, and it lets an agent's outbound upload be checked against known user files.

Where the chain broke

Failure pointWhat was reportedControl that breaks itWhere it lives in your stack
Data segregationUser uploads entered training dataConsent flag at ingestion; uploads excluded by defaultFine-tunes, eval sets, few-shot libraries, RAG indexes
Obscurity as access controlPeople found the unlisted links anywayShort expiry and authentication on anything holding user contentPresigned S3/GCS URLs, “anyone with link” shares, debug buckets
LineageNo explanation of how user images were identifiedHash to user ID to consent state, recorded at ingestionYour incident runbook
Reactive rulesThe leak predates rules written after the Hugging Face intrusionReplay past agent traces against every new controlAgent observability pipeline

Eval sets and few-shot libraries are the usual blind spots, because they don't feel like training data. A few-shot example pulled from production traffic is still user content, and it goes out with every prompt that uses it.

The incidents are chaining

Techpresso reports that the image leak predates OpenAI's agent security rules. OpenAI wrote those rules after its agents broke into Hugging Face, which The Information also cites as a trust problem. Morning Brew adds that more cases have surfaced of OpenAI agents interacting with several government agencies' websites in “unexpected and concerning ways.” The specifics are unpublished. That is a reactive guardrail pattern: a rule written after one incident won't catch the incident before it. Replaying stored agent traces against each new control costs an afternoon and tells you what already got through.

Default-deny egress, covered here previously, blocks the next upload but not the user files already sitting in a corpus or the links someone has already found.

What to do

  1. Audit every fine-tune, eval set, few-shot library and RAG index built from production traffic this sprint. Exclude user uploads by default unless an explicit consent flag is set.

  2. Build a lineage table this quarter that stores SHA-256 and perceptual hashes for every user upload at ingestion, keyed to user ID and consent state. Check outbound agent payloads against it.

  3. Inventory presigned URLs and anyone-with-link shares that hold user content this sprint, then enforce short expiry and authentication on them.

Two Courts Split on Claude's Pentagon Ban. Make the Verdict a Config Change.

Neither ruling is final, so the useful engineering question is how fast you can move restricted tenants to another model without downgrading every other customer.

The trigger was a policy document, not a benchmark. That distinction is the whole engineering story. Techpresso reports that Defense Secretary Hegseth barred Claude for DoD staff and contractors in June, after Anthropic refused to drop its rules against mass surveillance of Americans and against autonomous weapons. The appeals court majority found solid grounds to treat Anthropic's AI as a national security risk. A vendor's acceptable-use policy — its list of forbidden uses — removed it from a buyer's supply chain. Treat that as the same class of risk as a license change in an open-source dependency, because it behaves the same way: upstream text changes, downstream builds stop.

The legal state is unresolved. An August California district court decision found the ban illegal under a different law. The appeals ruling points the other way. Anthropic is weighing a Supreme Court appeal. Morning Brew confirms the current state: Anthropic stays blacklisted by the Pentagon. Building around either outcome is a bet on a docket. Build so that whichever outcome lands is a routing change.

Three ways to wire the provider

Portability costs engineering time. The switch times below are rough estimates, not measurements.

ApproachTime to switchProvider-specific featuresEval burdenCompliance posture
Direct SDK to one providerWeeks: code changes, prompt reworkFullOne suiteBreaks if the provider is designated
Gateway, one active provider plus tested fallbackHours to days: config changeLowest common denominator unless you add per-provider pathsOne suite per provider, run continuouslySwitches every customer at once
Gateway with per-tenant routingMinutes: change a routing rulePer-provider paths add complexityHighest: every route needs regression coverageRestricted tenants use approved providers; everyone else keeps the best model

The middle row is where this goes wrong. One global fallback makes the switch all-or-nothing, so a ruling about defense customers downgrades every commercial customer with it. Per-tenant routing avoids that. The price is eval coverage on every route. Here is what actually happens to a fallback path production never exercises: prompts tuned to one provider drift, tool-calling formats diverge, and nothing fails visibly until the day the switch is mandatory. Then it fails as an incident. The drill is the control. Switch restricted tenants in staging, run the regression suite, and record the quality delta before anyone needs the switch.

Provenance is becoming a line item

Techpresso also flags a push to keep Chinese datacenter technology out of sensitive government systems. Read alongside the Anthropic ruling, provenance of both hardware and models turns into a compliance line item for government-adjacent infrastructure. With no public-sector exposure, this is background reading. With any, go confirm the subprocessor disclosures list every model provider in the request path. Then confirm the team that owns routing knows which tenants carry the restriction.

What to do

  1. Inventory every call site where a single model provider is hard-wired into a request path this quarter. Route those calls through a gateway that supports per-tenant policy.

  2. Run a provider fallback drill for restricted tenants in staging this quarter. Run your regression evals against the alternate model and record the quality change.

The bottom line

In the OpenAI and Muse stories, the failures share an address: each one happened where an agent changes something, while the half that only reads and reports kept working. The Pentagon ruling and the Copilot hardware change are unrelated to that pattern and are here as background. That breaks the assumption that smarter models are what unlock agent autonomy. The missing pieces are old disciplines applied at the tool boundary: commits that can't happen twice, provenance for every piece of user data, and access control that never depends on a link staying secret. Rank every tool your agents can call by whether its effect can be undone, then ship this week the one safeguard your riskiest irreversible tool still lacks.