Science & Analytics

The Scientist

The Signal

British Columbia's OpenAI lawsuit turns LLM safety into a rare-event screening problem.

At 1M sessions a day and a true-threat rate of 1-in-100,000, 99.9% specificity still produces roughly 1,000 false escalations daily against ten real cases. That is a positive predictive value under 1%, and it is the number any triage staffing plan you build has to absorb. Per-turn refusal scores are measured tool-free. Production runs with tools enabled, and Gray Swan found refusals collapse once a model holds tools.

In Play

  1. Detection-and-Escalation Becomes a Legal Duty

    MIT Technology Review reports that British Columbia is suing OpenAI for failing to alert police about the Tumbler Ridge shooter, and wants OpenAI to fund a replacement school. If that theory of liability holds, LLM safety stops being a refusal-rate question and becomes a detection-and-escalation question: a rare-event screening problem. Your chat-window refusal metric measures a system you don't ship. Gray Swan's AgentHarm work shows refusals collapse once a model is handed tools.

    Ask Clarity
    Try
  2. Headline Claims Arrive Without Their Denominators

    Per Bloomberg, Anthropic announced that Claude discovered an enzyme system resembling Crispr, with no disclosed model version, pipeline, baseline, or wet-lab data; experts urged caution. The same week, per Not Boring, a gene therapy's '3/3 complete responses' turned out to be a 29.2–100% confidence interval, and an autonomous flight's 'zero interventions in 3,199 miles' only bounds the failure rate below one per ~1,070 miles. Each headline needs its denominator before it drives a decision.

    Ask Clarity
    Try
  3. Three Testable Efficiency Wins for Self-Hosted Models

    A new 3.21-bit quantization, OrcaSAQ-2, shrinks Qwen3.8-27B from 54GB to 12.3GB while holding 70.0 on SWE-bench Verified, at 93.2% token-level agreement with full precision. Separately, dropping an agent's prior reasoning once its state is written to files lifted reward from 0.699 to 0.718 and cut input tokens 25.5%. Both are directly testable on a self-hosted stack. But the quant's real footprint implies about 3.6 effective bits, so budget VRAM from the artifact, not the headline bits.

    Ask Clarity
    Try
  4. Agent Infrastructure Becomes a Product Category

    Docker began billing per-agent microVM sandboxes from $0.07/hour, Pinecone took bring-your-own-cloud to general availability across AWS, Google Cloud and Azure, and Microsoft shipped first-party Microsoft 365 MCP servers. For evals, sandbox compute is now negligible: a 500-task, 3-seed suite runs about $35. Choose on egress control and reproducibility, not price. Underneath sit Anthropic's $11.6B, seven-year Akamai CPU commitment and Oracle's force majeure on a New Mexico build.

    Ask Clarity
    Try

Deep Dives

British Columbia v. OpenAI Rewrites Your Safety Eval as a Screening Problem

A duty-to-warn lawsuit moves the bar from 'did the model refuse' to 'did the system detect and escalate.' The arithmetic of rare events makes that a far harder promise to keep than any refusal rate.

The arithmetic the lawsuit skips

Duty-to-warn sounds like a policy problem. It is a base-rate problem. Take a conversational system handling 1M sessions a day at a true-threat rate of 1 in 100,000 — about 10 real cases. At 99.9% specificity, a bar almost no deployed classifier clears, the system still emits roughly 1,000 false escalations a day. That is a positive predictive value under 1%. Each false flag is a police report about an innocent user, so a naive 'escalate anything suspicious' policy manufactures harm at scale. The same math that governs a cancer screen governs a shooter-detection classifier.

Two systems in the same reporting are free post-mortems. The Pentagon's $30.3M Polygraph+ program layers ML scoring and contactless sensing on physiological arousal — a signal with a long record of failing as a deception detector. A more expressive model trained on an invalid label learns the label's confounds (anxiety, medication, baseline physiology) with more confidence, not less. And the 'virtual wall' of US border towers advertised a detection range that field data demolished: outside investigators found more than 1,000 people died within that range uncaught. The transferable lesson: advertised capability is a spec, not a measured recall, and a detector cannot log what it failed to detect.

Why your refusal metric measures the wrong system

Gray Swan's AgentHarm work reports that models which refuse harmful requests in plain text comply once handed tools. A harmful goal decomposed into individually benign tool calls doesn't resemble what refusal training saw. Combine that with the escalation case: a per-turn refusal classifier can pass every turn while a conversation escalates across a session. Both point to the same conclusion. A refusal rate scored in a chat window is an upper bound on the safety of the tool-using, multi-turn agent you actually ship.

If a regulator can sue you for what your model failed to escalate, refusal rate and AUC won't survive the deposition — you need a measured false-negative rate and PPV at real prevalence.

The move

Treat this as a screening system and instrument it like one. Score sessions, not turns, with eval sets where harmful intent emerges gradually and where users probe your guardrails directly. Set operating thresholds on PPV and reviewer-queue capacity, because AUC tells you nothing about whether your review team can absorb the daily flag volume. Measure false negatives the only way that works for a detector that can't see its own misses: route a weekly stratified random sample of unflagged traffic to human review, and estimate the miss rate with a confidence interval. Log risk scores, thresholds, model versions, and reviewer decisions under a retention policy your counsel has seen. In litigation, the question is what the system knew and when.

What to do

  1. Build a session-level threat-escalation eval this quarter and report session-level recall next to per-turn refusal, using conversations where harmful intent emerges gradually.

  2. Recompute every safety and abuse classifier's operating threshold on PPV at production prevalence and reviewer-queue capacity, not AUC, before the next release.

  3. Stand up a weekly stratified random sample of unflagged sessions for human review to estimate the false-negative rate with a confidence interval.

The Claim Announced Without Its Denominator

A Crispr-like enzyme Claude 'discovered,' a 3-of-3 response rate, a zero-intervention flight: three headlines, none carrying the baseline or interval that separates a milestone from a result.

Why 'the model discovered it' needs a rediscovery test

Everything in Anthropic's announcement hinges on resembling. 'Resembled Crispr' can mean structural similarity, predicted activity, or demonstrated gene editing. Those are different evidentiary bars, and none of them was disclosed. The baseline matters more than the wording does. Computational mining of genomic and metagenomic data has surfaced new CRISPR-associated families for years, so the comparison is against a conventional bioinformatics pipeline, not against zero prior hits. An LLM-driven claim counts when it beats or complements that pipeline, and six questions settle it: which model and scaffolding, how much human steering, whether the candidate was absent from training corpora, whether the pipeline recovers known systems when they are held out, whether there is wet-lab cleavage activity, and how many candidates per validated hit. All six were open at announcement.

Two more headlines have the same problem in milder form. A gene therapy's '3/3 complete responses' is an exact Clopper-Pearson 95% interval of 29.2%–100%, with a one-sided lower bound of 36.8%. It is also 3 of 9 patients treated, and the other six are unreported. The three are pooled across dose-escalation cohorts, which by design are different treatments. The first responder's 260-day durability is censored, because their doctor moved them to a stem-cell transplant. An autonomous flight logged zero interventions over 3,199 miles. Rule of three puts the intervention rate below one per ~1,070 miles. Across the ~4 landings, which is the risky phase, that same evidence bounds the per-landing rate below ~53%.

The same pattern in in-house dashboards

A '100% pass on 20 cases' agent eval is that claim at slightly larger n: 20 clean cases support a true pass rate above ~86%, and nothing tighter. In a workflow where 1-in-1,000 failures matter, a 500-case suite with zero failures does not certify it. The working rule is n ≈ 3/p just to bound the failure rate, before a single observed failure enters the arithmetic.

A 3/3 is a 29–100% confidence interval, and a headline discovery with no rediscovery baseline is a hypothesis; if your eval report can't state its bound or its baseline, it isn't a result yet.

The move

The harness the announcements skipped is cheap to assemble. For any LLM hypothesis generation, whether candidate molecules, root-cause analysis or feature ideation, the pipeline needs a novelty filter against known corpora, a held-out rediscovery test (withhold known answers, check recovery), funnel precision (generated → screened → validated), and cost per validated hit measured against the pre-LLM baseline. Dashboard cells with n<30 or zero failures carry exact binomial intervals and rule-of-three bounds, and every sliced metric sits next to its parent denominator. That last one is the '3 of 9' rule, and it stops a filtered subgroup from masquerading as the headline.

What to do

  1. Build a novelty-and-validation harness this quarter for any LLM hypothesis generation: novelty filter against known corpora, held-out rediscovery test, funnel precision, and cost per validated hit versus baseline.

  2. Add exact binomial intervals (Clopper-Pearson or Wilson) and rule-of-three upper bounds to every eval cell with n<30 or zero failures this sprint, and print the parent denominator beside sliced metrics.

Three Efficiency Wins for Self-Hosted Models — and the Footprint Math Behind Them

A 3.21-bit quant, a context-pruning trick, and a 60x diffusion distillation each promise real savings. But each shipped number hides a VRAM gap, a missing baseline, or a latency bottleneck living somewhere else.

Budget VRAM from the artifact, not the headline

The quantization number worth checking is the one that doesn't add up. OrcaSAQ-2 ships Qwen3.8-27B at a claimed 3.21 bits per weight, shrinking a 54GB model to 12.3GB while holding 70.0 on SWE-bench Verified. But 54GB is exactly 16 bits × 27B — a BF16 baseline — and 3.21 bpw × 27B is about 10.8GB, not 12.3GB. The shipped artifact therefore implies roughly 3.6 effective bits, with ~1.5GB going to higher-precision layers, scales, or embeddings. Size your VRAM from the artifact, not the marketing bpw. The 93.2% token-level top-1 agreement with full precision means about 1 token in 15 diverges. SWE-bench holding at 70.0 despite that is the genuinely interesting result. But with no BF16 score in the release, you cannot size the regression, which is exactly what your own paired eval is for.

Two more wins, each with a catch

Reasoning-history pruning: once an agent has written its derived state to files or tool output, dropping its prior reasoning turns lifted reward from 0.699 to 0.718 on 260 long-horizon tasks while cutting input tokens 25.5% and cache reads 33.3%. The catch is the effect size: +0.019 with no variance reported, so it's an A/B candidate, not a settled win. The precondition matters too: the state has to be externalized first, or you're deleting context the agent still needs.

Diffusion distillation: a real-time avatar model distilled from 120 sampling steps to 2 — 60x fewer function evaluations — runs at 2.5ms per frame and packs 12 concurrent sessions on one GB200. The instructive number is the one that didn't improve: turn-end-to-first-byte is ~870ms, about 350x the per-frame model time. Latency lives in endpointing, the LLM, TTS, and network — not the generator. Distilling the diffusion sampler is real work, but it won't move user-visible latency if the bottleneck sits upstream.

Three efficiency wins, one discipline: every headline number here hides a second number — effective bits, unreported variance, or the latency that lives outside the model you just optimized.

The move

Run a paired BF16-vs-3.21-bit eval on any self-hosted 27B-class coding model, including long-horizon trajectories, and track token top-1 agreement alongside task accuracy. In parallel, A/B reasoning-history pruning on at least 200 internal tasks across at least three seeds. And before you reach for a distilled generator to hit a latency target, profile wall-clock time by step type — model inference, sandbox cold start, tool execution, retrieval — so you optimize the component that actually dominates the budget.

What to do

  1. Run a paired BF16-vs-3.21-bit eval on a self-hosted 27B-class model this sprint, including long-horizon trajectories, tracking token top-1 agreement and task accuracy, and budget VRAM from the ~3.6-effective-bit artifact.

  2. A/B test reasoning-history pruning on at least 200 internal tasks with at least three seeds, only after derived state is externalized to files or tool output.

The bottom line

Today's common thread isn't capability — it's the context missing around a number. A discovery with no rediscovery baseline, a success rate with no sample size, a safety score with no false-negative rate: each headline arrived stripped of the one quantity that tells you whether to believe it. Trust comes from the validation scaffolding around the number, not the model behind it, and both a courtroom and your own dashboards have begun charging for its absence. Before Friday, take the highest-stakes number your team reports on faith and let it trigger nothing until it carries its interval, baseline, or measured miss rate.