Science & Analytics

The Scientist

The Signal

88% of 10,616 AWS keys leaked over four years still authenticated on re-check.

The median key sat five years unrotated, and 8.2% of the live ones carry full admin. That pins down the asymmetry your secret detector's threshold has been guessing at: a miss stays exploitable for years, while a false positive costs one triage ticket. F1 scores those two errors as equals, which is the part the leaderboard number never tells you.

In Play

  1. Cost Matrices Land for High-Stakes Classifiers

    Write the cost matrix for your highest-stakes decision surface this week. GitHub's secret-scanning playbook ships without a single precision or recall figure; Truffle Security's re-verified leaked AWS keys supply the loss function it was missing. First deep dive has the arithmetic.

  2. Open-Weight Fine-Tunes and Where Model Spend Actually Goes

    Bridgewater's open-weight fine-tune and Ramp's card-spend index both point at routing and fine-tunes, not frontier checkpoints, as the cost lever. Vercel's 62% open-weight token share is a single-day maximum, not a trend. Second deep dive prices the breakeven.

  3. Agent Evaluations Are Scored at the Wrong Unit

    Five agent-security papers converge on one structural flaw: guardrails reset with each task, so an attack distributed across many steps never trips a threshold. All five are unnamed and unlinked, so treat the numbers as preprint-grade signal, not results.

  4. Android 17 Schedules a Covariate Shift Into Your Features

    Android 17 ships OS-wide Encrypted Client Hello, per The Hacker News. Any model consuming SNI, hostname tokens or TLS-handshake metadata loses those features structurally as the fleet upgrades: recall falls quietly on the newest-OS cohort while aggregate drift monitors stay in band. No rollout curve was published, so the slope is unmodellable today.

  5. Two Label-Validity Failures in Everyday Pipelines

    A lifecycle model scored a reader as churned the week he died, because mortality is an unobserved competing risk inside every engagement-based churn label. Separately, agent-dominated documentation traffic breaks the human-behaviour assumptions built into session definitions and A/B randomisation. Third deep dive covers both fixes.

Deep Dives

  1. Truffle's Base Rates Make the Threshold Decision Arithmetic

    GitHub shipped the operating discipline for LLM classifiers with no numbers attached, and a separate re-verification of four years of leaked credentials supplies the loss function it was missing.

    Price the miss before tuning the threshold Start with the survival curve. Truffle Security re-verified 10,616 AWS keys leaked publicly between August 2022 and August 2026. 88% still authenticate. Of the roughly 9,342 live keys, 768, about 8.2% , carry…

    3 action items

  2. The 14x Saving Is Priced Before Anyone Buys Labels

    Bridgewater's fine-tune beat every frontier model it tested, but the asset doing the work is a dataset most teams lack — here is the volume threshold that decides who can copy the result.

    The breakeven, written out The decision rule is a breakeven volume rather than a cost ratio: V ≈ L / (C_frontier − C_finetuned) , with L the total labeling plus training spend and the C terms per-token inference costs. A…

    3 action items

  3. Your Churn Label Contains a Hazard It Can Never Observe

    Mortality is one of four processes hiding inside a single binary target, and the fix is a formulation change plus an external feed rather than a bigger model or better features.

    Four hazards wearing one coat A binary churn target collapses at least four generating processes into one event: voluntary disengagement, deliverability failure, identity or account change, and death. The hazard shapes differ, and so do the correct interventions. What breaks…

    3 action items

The edition continues

Take the signal into the room.

Sign up or log in to read all 3 deep dives in full, plus the final take.

Read the full edition

Continue with LinkedIn