Science & Analytics
The Scientist
62 live API keys sat in public agent traces that no redaction pipeline can read.
The envelope authenticates model name and version. It does not authenticate the issuing account or the conversation, so any cheaper same-family model decrypts and replays it. Scrubber recall was measured against plaintext, and against ciphertext that recall is zero, which is the only slice that matters for the bundles you have already published. Those stay unredactable for good. Familiar shape, secrets sitting in something nobody thought of as a secret store, except this time the cleanup pass cannot see the thing it was hired to find.
In Play
Encrypted Reasoning Blocks Behave Like Bearer Tokens
Encrypted reasoning blocks replay into cheap sibling models and print the flagship's hidden trace — 62 live API keys recovered from public agent traces you cannot redact. Every item below breaks at the same joint: the artifact you measure has quietly diverged from the quantity you infer from it. Researchers at MATS, ELLIS Tübingen and the Max Planck Institute for Intelligent Systems replayed the blocks Anthropic, OpenAI and Google hand back to API clients into cheaper same-family models, per ByteByteGo. Plaintext redaction has zero recall on ciphertext, so any trace bundle you have already published is unredactable. Gateway owner, this week: reject envelopes whose header model and version do not match the model being called.
Ask ClarityToken-Savings Claims With No Pass Rate Column
Atlassian's ~50% token cut with Codex or Claude Code, its Code Context claim of 44% more accuracy on 48% fewer tokens, and Kiro's roughly 82% lower cost with GPT-5.6 Terra all ship without a task-success rate, per Applied AI and TLDR IT. A token cut with no pass rate beside it is indistinguishable from context truncation. At 50 cases the noise floor is ±13pp — wide enough to manufacture the win. Eval owner, before the next cheap-configuration swap: pre-register the non-inferiority margin and the held-out split.
Ask ClarityDrift Telemetry Drops Data Exactly at the Spikes
Telemetry loss is load-correlated, so p95 latency, population drift statistics and A/B guardrail metrics are all computed on a survivorship sample. ClickHouse disclosed that in-memory collector queues and local write-ahead logs could not handle OpenTelemetry ingestion under database backpressure, and rebuilt LogHouse around priority failover with blob storage as durable overflow — now running at 50 million events per second, per TLDR IT. Roughly $5.35B of observability assets changed hands over months, including Dynatrace's $915M purchase of Arize. Platform owner, this quarter: prove your collector spills to durable object storage under sink backpressure before trusting another drift number.
Ask ClarityDistillation Has a Law (Apple, Early 2025), and It Says Pick Teachers by Cross-Entropy
Apple's Distillation Scaling Laws (Busbridge et al., early 2025) swept students from 143M to 12.6B parameters over budgets up to 512B tokens. Student loss tracks the teacher's cross-entropy on the distillation distribution, not teacher size — so a stronger teacher can yield a worse student. Production KD's 70B-plus teachers sit far outside the sweep, so the shape may transfer but the constants will not, per TheSequence. Model team, before the next KD run: select the teacher by measured cross-entropy on your own distillation set, not by parameter count.
Ask ClarityThe Label Source Is the Confounder, Not the Model
Palo Alto Networks found only 12 of more than 405 AI-enabled malware samples in real-world infections; the rest live in sandboxes and VirusTotal as research and red-team artifacts, per Risky Business. Meanwhile a Treasury action implies roughly 30 months of undetected Iranian-linked exfiltration across five U.S. sectors — a window an AUC-scored detector treats as acceptable recall, per CyberScoop. The sampling frame and the scoring metric decide the number before any model does. Detection team, this quarter: state your prevalence assumption and dwell-time recall next to every AUC you report.
Ask Clarity
Deep Dives
- ●
Your Trace Store Is Holding Secrets Nobody Can Read
The cryptography held and the threat model didn't — and the same missing field means your faithfulness numbers are scoring a summary rather than the reasoning that produced the answer.
The missing field The envelope is a valid AEAD construction and its integrity guarantee held. What is absent is a binding between the block and the context that produced it . The authenticated fields cover block content plus model name…
- ●
Four Efficiency Claims, Not One Pass Rate Between Them
Automated optimizers now search the same fifty-case suites these vendor numbers came from, and at that sample size the noise floor is wide enough to manufacture a ten-point win.
Do the arithmetic on the suite you actually own Microsoft's Agent Lightning v1.0.1 installs into Claude Code, Codex, Copilot or Cursor and reworks prompts, tool definitions, workflow topology, model assignment and reasoning effort against a scored benchmark, per Unwind AI.…
- ●
The Pipeline Behind Your Drift Numbers Fails Under Load — and Changed Owners Over Months
A streaming vendor rebuilt its own ingestion because default collectors drop data at exactly the moments worth measuring, and the layer storing those traces absorbed billions in acquisitions over months.
The missingness is informative, and no imputation fixes it The collector's in-memory queue overflows when the sink stalls under load. The observations it drops are the ones generated during the windows worth measuring: traffic spikes, cache-miss storms, retry cascades. The…
The edition continues
Take the signal into the room.
Sign up or log in to read all 3 deep dives in full, plus the final take.
Read the full editionContinue with LinkedIn