Science & Analytics
The Scientist
OpenAI's model decrypted its own reasoning when distillers pasted it into a new chat.
No encryption was broken and no stored chats were touched, because the encrypted blob was never bound to the conversation that issued it. The campaign spread 16,000 requests across 4,000+ accounts in two days, about four each, too thin for any per-account limit or abuse classifier to see, including the ones you are probably running. Those detectors measure per-account rates, and this campaign kept every account's rate unremarkable.
In Play
ML serving code treats bytes as trusted
SANS @RISK reports that SGLang's multimodal runtime passes unauthenticated ZeroMQ messages straight into pickle.loads(). The flaw is CVE-2026-93088, rated CVSS 9.8. CSO separately reports that just selecting a model in Unsloth Studio could run Python from that model's repo. Both bugs sit on hosts that hold GPU access, weights and cloud tokens. Port-scan those hosts from off-host today to confirm nothing is reachable, then rotate tokens and tighten how models get loaded this week.
Ask ClarityDistillation hid under per-account limits
The Information's account of OpenAI's forensics shows the July distillation campaign peaked at 16,000 requests from 4,000+ accounts over July 24–25. That works out to about 4 requests per account. Per-account rate limits and abuse classifiers cannot see traffic that thin; only clustering across accounts can. If you serve any model, including an embedding or ranking endpoint, you are probably watching the wrong unit.
Ask ClarityMetric definitions beat model swaps
TLDR Data reports that adding a model-built semantic layer lifted LLM analytics accuracy from 39% to 91%. The errors that remained came from hidden business conventions, such as which accounts count as customers. That gap is likely larger than the spread between frontier models on the same task. Google's PageBreak and a proposal to have LLMs write linters point the same way: deterministic structure around the model sets its precision.
Ask ClaritySpeed and memory get pricier as tokens deflate
OpenAI's Ultrafast tier charges 6x on both input and output tokens but only speeds up output. By Simplifying AI's derivation, each hour of waiting you buy back costs about $54 plus $10.8 times your input:output ratio. Separately, The Information reports HBM sold out through next year and Nvidia systems up about 17%. Devansh puts the fall in price for a fixed capability tier at ~940x in three years, so your token mix and memory footprint now drive cost more than list price does.
Ask Clarity
Deep Dives
- ●
SGLang, Unsloth and MCP: the files your ML hosts load are executables
Four separate flaws share one design error, treating model repos, RPC messages and tool servers as inert data, and each one lands on machines holding your weights and tokens.
Four bugs, one design error Component What goes wrong Severity First move SGLang multimodal runtime Unauthenticated ZeroMQ ROUTER socket feeds pickle.loads() CVSS 9.8 (CVE-2026-93088) Firewall the socket; upgrade per advisory Unsloth Studio model picker Selecting a model can run its…
3 action items
- ●
Thin traffic, thick signal: why per-account defenses missed the distillation campaign
The attack exploited a protocol bug, not a model weakness, and it was visible only across accounts, where most serving teams collect no telemetry at all.
A replayed blob, decrypted on request OpenAI's forensics, as The Information describes them, show operators copying encrypted reasoning from one conversation into another and asking the model to decrypt and transcribe it. CyberScoop adds that no encryption was broken and…
3 action items
- ●
Your text-to-SQL accuracy lives in your metric definitions
Business rules, not a better model, produced a 52-point gain, and the same kind of cheap deterministic structure is how Google got a hallucination-prone agent to near-zero false positives.
It's a label-definition problem, not a reasoning problem The useful part of TLDR Data's result is the failure analysis. After the semantic layer went in, the errors that remained were not arithmetic. They were hidden business conventions , such as…
3 action items
The edition continues
Take the signal into the room.
Sign up or log in to read all 3 deep dives in full, plus the final take.
Read the full editionContinue with LinkedIn