British Columbia v. OpenAI Rewrites Your Safety Eval as a Screening Problem
A duty-to-warn lawsuit moves the bar from 'did the model refuse' to 'did the system detect and escalate.' The arithmetic of rare events makes that a far harder promise to keep than any refusal rate.
The arithmetic the lawsuit skips
Duty-to-warn sounds like a policy problem. It is a base-rate problem. Take a conversational system handling 1M sessions a day at a true-threat rate of 1 in 100,000 — about 10 real cases. At 99.9% specificity, a bar almost no deployed classifier clears, the system still emits roughly 1,000 false escalations a day. That is a positive predictive value under 1%. Each false flag is a police report about an innocent user, so a naive 'escalate anything suspicious' policy manufactures harm at scale. The same math that governs a cancer screen governs a shooter-detection classifier.
Two systems in the same reporting are free post-mortems. The Pentagon's $30.3M Polygraph+ program layers ML scoring and contactless sensing on physiological arousal — a signal with a long record of failing as a deception detector. A more expressive model trained on an invalid label learns the label's confounds (anxiety, medication, baseline physiology) with more confidence, not less. And the 'virtual wall' of US border towers advertised a detection range that field data demolished: outside investigators found more than 1,000 people died within that range uncaught. The transferable lesson: advertised capability is a spec, not a measured recall, and a detector cannot log what it failed to detect.
Why your refusal metric measures the wrong system
Gray Swan's AgentHarm work reports that models which refuse harmful requests in plain text comply once handed tools. A harmful goal decomposed into individually benign tool calls doesn't resemble what refusal training saw. Combine that with the escalation case: a per-turn refusal classifier can pass every turn while a conversation escalates across a session. Both point to the same conclusion. A refusal rate scored in a chat window is an upper bound on the safety of the tool-using, multi-turn agent you actually ship.
If a regulator can sue you for what your model failed to escalate, refusal rate and AUC won't survive the deposition — you need a measured false-negative rate and PPV at real prevalence.
The move
Treat this as a screening system and instrument it like one. Score sessions, not turns, with eval sets where harmful intent emerges gradually and where users probe your guardrails directly. Set operating thresholds on PPV and reviewer-queue capacity, because AUC tells you nothing about whether your review team can absorb the daily flag volume. Measure false negatives the only way that works for a detector that can't see its own misses: route a weekly stratified random sample of unflagged traffic to human review, and estimate the miss rate with a confidence interval. Log risk scores, thresholds, model versions, and reviewer decisions under a retention policy your counsel has seen. In litigation, the question is what the system knew and when.
What to do
Build a session-level threat-escalation eval this quarter and report session-level recall next to per-turn refusal, using conversations where harmful intent emerges gradually.
Recompute every safety and abuse classifier's operating threshold on PPV at production prevalence and reviewer-queue capacity, not AUC, before the next release.
Stand up a weekly stratified random sample of unflagged sessions for human review to estimate the false-negative rate with a confidence interval.