Output Quality Is Now a Policy Decision, Not a Specification
Three federal agencies asked American labs to serve worse answers to suspected distillers without telling them, and the traffic profile they described is an ordinary enterprise agent pipeline.
The mechanic deserves a second read from procurement. The advisory tells labs to reduce reasoning depth, present correct information with different reasoning, introduce stylistic inconsistencies, and vary the degradation across requests "to complicate response quality evaluations". It also tells them not to inform the affected users. Safety researchers and third-party evaluators get an explicit carve-out and should be told.
The carve-out is the tell. Exempting evaluators concedes that degradation defeats measurement, which means a two-tier reliability regime now exists in a market where every SLA, eval score and cost-per-task model assumes a stable endpoint. Sophisticated buyers will negotiate into the top tier. Everyone else keeps paying specification prices for a discretionary output and never sees the difference, because the recommended pattern is built to survive a single-shot test.
The second front is provenance, and it arrives as a document request
The same advisory accuses Moonshot AI of distilling 17 American models, including Anthropic's Claude Fable 5, released only months earlier. Moonshot declined to comment. Its Kimi K3 is described as a genuine global hit, which is the reason the advisory exists at all. This is not litigation. It is a procurement artifact that any enterprise security reviewer, government buyer or acquirer can cite from next week, and exposure counts when it is indirect: through a multi-model router, a fine-tuned open-weight derivative, or a SaaS vendor's undisclosed inference stack.
Confidentiality moved in the same cycle. OpenAI, disputing an attribution claim, conceded it "cannot rule out that de-identified data derived from their usage of our products helped improve our models." Read that alongside model vendors pushing into finance workflows and payments, and one exposure has two faces. The supplier can vary what it delivers, and may learn from what it receives.
Where the reporting converges, and where it splits
Four independent accounts agree on the facts of the advisory and on the direction. Model supply chains are now a documented, disclosable part of enterprise risk. They split on the first move. One line of analysis says contract first: a non-degradation and disclosure clause. Another says documentation first: a model bill of materials naming every model, router and derivative in production within ten days. A third argues for a 30-day lineage sweep across production, fine-tuning and evaluation stacks, on the view that country-of-origin becomes a procurement gate before enforcement arrives.
They are sequencing the same three artifacts, and the order that holds is cheapest-first. The clause costs a letter. The bill of materials is what the first customer question demands, and producing it under deal pressure turns a memo into a discovery project. The measurement harness comes third and matters most, because longitudinal telemetry is the only evidence base that would ever support a contractual claim.
A vendor's refusal to put non-degradation in writing is itself the intelligence, and it belongs in the risk register rather than the follow-up folder.
A reasonable skeptic says no lab will ever sign this. Some won't. The negotiation is still cheapest now, before the practice normalizes and vendor legal positions harden. The predictable failure is filing all of it with the security team. This is supply-chain integrity with revenue, contract and disclosure consequences, and it needs an owner who sits at the executive table.
What to do
Send every frontier model vendor a written demand this week for a non-degradation and disclosure clause: no undisclosed model substitution, a notification SLA on quality-affecting changes, and confirmation your accounts are not classified as a distillation risk.
Publish a model bill of materials within ten days naming every model, router and fine-tuned derivative in production, flagging exposure to the six named firms including through vendors.
Fund a continuous canary evaluation harness this quarter: a fixed golden set run hourly against every vendor endpoint, tracking correctness, reasoning depth, latency and style drift with statistical alerting.