The Provenance Marker You Never Put In A PRD
Detection tooling is not public yet, and that lag is the only reason your disclosure copy can still be fixed quietly instead of during a customer escalation.
What the marker actually proves
The first customer to notice will not have read the announcement. They will get a hit somewhere downstream and open a ticket. What they hit is keyed sampling bias: the model nudges word choice into a statistically detectable pattern that a z-score test can recover afterward. AINews's read on the technical detail governs product decisions more than the announcement does. The signal degrades under paraphrasing or regeneration, false positives on naturally written text are unresolved, and no third-party detection tooling exists. Anthropic's own framing is that a hit means text may have been processed by Claude, not authored by it.
That combination is awkward in a specific way: weak evidence, strong obligation. Not solid enough to build a feature on. More than solid enough to produce a customer question with no good answer. Devshot's framing is the accurate one. A governance property arrived in the product with no PRD, no release note, and no customer note from the product side.
Where the sources agree, and where they split
Agreement is total on reach: applied at the model level, so it travels through the API, Claude Code, Cowork, and Claude hosted on AWS, Google Cloud and Microsoft Foundry, worldwide. The split is about what kind of problem this is. Simplifying AI ties it to EU transparency rules that apply globally, the cleanest Brussels-effect precedent yet, which retires "scope AI compliance work as EU-only" as a planning default. AI Breakfast reads it as competitive positioning: Anthropic pairs text watermarks with C2PA metadata on image outputs and converts provenance into an enterprise-safe sales line. AINews reports the backlash. A thread with 2,077 activity points on r/LocalLlama shows users citing Claude-linkable marking as a reason to prefer open weights.
Take all three, in that order. The compliance framing sets the deadline. The positioning framing says what provenance is worth once it has been disclosed, and the backlash says a segment of existing users will treat marked output as a defect and ask for an alternative model path. That segment shows up in retention data, not in a launch deck.
The feature to kill before someone specs it
Someone will propose surfacing detection to users. The Algorithmic Bridge's analysis of Substack's Pangram 4.0 integration is the cheapest available lesson. The detector reports a 0.0041% false positive rate against a 0.34% false negative rate, roughly 83 times more willing to miss a cheater than to accuse an innocent, and it is still brittle on short texts. Chunk-based processing can score a document 100% human overall while individual passages score 100% AI. Two more traps ship with that design. A pre-publish self-check hands evaders an unlimited adversarial-testing oracle against a production detector. An optional badge always collapses: a visible opt-out reads as guilt, an invisible one makes the badge worthless.
Any surface that accuses a user needs a signed-off false-positive budget in the PRD, not an accuracy number in the launch post.
Caveat worth holding: every Pangram figure is vendor-reported, and the author who published them says plainly that he does not believe the product is as good as claimed.
What changes on the product side
Three artifacts, none of them engineering work: the AI-disclosure page, the last three security-questionnaire responses, and a support macro explaining that the marker is non-dispositive in both directions. Write the macro before the ticket arrives. The first person to ask will be a customer whose published work got flagged somewhere downstream. Being discovered is categorically worse than disclosing.
What to do
Map every surface where Claude-generated text or images reach a customer — direct API, Bedrock, Vertex, Microsoft Foundry, Claude Code inside the dev loop — and rewrite the AI-disclosure page plus security-questionnaire answers as a near-term priority to state that outputs carry a machine-readable provenance signal.
Kill any roadmap item that treats a watermark hit as proof of AI authorship, and replace it this sprint with a document-level 'provenance signal present' state plus a minimum-length threshold below which no score renders.
Hand Legal and Support a one-page provenance position and a published help-center answer by the end of this sprint, covering what the signal is, what it does not prove, and which of your surfaces emit it.