The Base Rate Flipped: Human Text Is Now the Rare Class
Contamination is a detector problem before it is a data problem, and every synthetic-text classifier in your ingest path is calibrated on a prior that no longer exists.
Start with the classifier, not the corpus
Pew's estimate comes from scoring pages with Open Pangram, and the study publishes no precision, no recall, and no human-labeled ground truth. Two properties of that measurement matter more than the headline number. First, the label collapses full generation and light copy-editing into a single bucket, so 35% is an upper bound on a fuzzy construct rather than an estimate of synthetic share. Second, the discriminating stylometric features, em-dash frequency and Oxford commas among them, rose over the study window. That is what you would expect when human writers absorb model style. The detector's features are being contaminated by the phenomenon it measures.
MIT Technology Review's item makes the arithmetic worse in a useful way. If a Nature-cited study finds AI fingerprints in 90% of biomedical papers, then inside that domain human authorship is the rare class. Positive predictive value collapses toward the prior. A positive prediction carries almost no information because nearly everything is positive. Any classifier tuned on a roughly balanced dev set in 2023 now sits at the wrong point on its own PR curve, and reporting accuracy on a balanced holdout does not just fail to inform. It misleads.
Where the numbers agree, and where they don't
| Figure | Population | What it supports |
|---|---|---|
| 10% AI-authorship | 10,000 random English pages, July 2026 | Nothing about newly published content — diluted by pre-2022 pages |
| Over one-third / 35% | Pages published after Nov 2022 | A contamination prior for any recent crawl |
| ~10x gradient (.com vs .edu/.gov) | TLD strata within the same sample | Ingestion weighting you can ship this week |
| 90% of biomedical papers | Stylometric prevalence, field-specific | Detector recalibration; not a fabrication rate |
The three reports agree on direction and disagree only where denominators differ, which is the cleanest kind of disagreement to have. The TLD gradient is the operational residue. Commercial domains carry roughly ten times the flagged rate of .edu/.gov, which sits near 1%, with .org around 5%. That supports a defensible weighting scheme and a candidate source for a frozen gold set that does not require trusting a detector at all.
The two systems this lands on
The reflex reading is training-data contamination and model collapse. The nearer-term damage is to measurement. If gold references, preference pairs, or "human-written" baselines were scraped after 2022, part of the measured human-versus-model gap has quietly become model-versus-model. Win rates inflate and regressions go invisible. That failure mode never shows up as a red dashboard.
The second system is drift monitors. Anything baselined on pre-2023 crawl now compares today's inputs against a distribution that no longer exists, so some fraction of the drift firing alerts is corpus composition change rather than user behavior change. The thing the alert doesn't tell you is which. RAG indices decay on the same mechanism: retrieval over an increasingly synthetic corpus degrades groundedness without moving any metric currently on the board.
The only contamination estimate you can produce without trusting a detector whose operating point you no longer understand is a timestamp.
The one thing not to build is a hard filter. With false-positive rate unpublished and features drifting, deletion bakes in a bias you cannot measure or reverse. Persist the authorship score as a retained feature and log its distribution per source domain instead.
What to do
Date-partition every scraped eval set at 2022-11-30 this sprint, re-run your primary model comparison on the pre-cutoff slice, and report the win-rate delta as your contamination estimate.
Recompute prevalence-adjusted PPV/NPV and a full PR curve for any AI-text classifier in ingest or trust-and-safety before it filters another document, then demote it to a retained feature.
Add crawl-date and TLD as first-class ingestion dimensions this quarter and publish a synthetic-share estimate per index shard on every RAG build.