Science & Analytics

The Scientist

The Signal

62 live API keys sat in public agent traces that no redaction pipeline can read.

The envelope authenticates model name and version. It does not authenticate the issuing account or the conversation, so any cheaper same-family model decrypts and replays it. Scrubber recall was measured against plaintext, and against ciphertext that recall is zero, which is the only slice that matters for the bundles you have already published. Those stay unredactable for good. Familiar shape, secrets sitting in something nobody thought of as a secret store, except this time the cleanup pass cannot see the thing it was hired to find.

In Play

  1. Encrypted Reasoning Blocks Behave Like Bearer Tokens

    Encrypted reasoning blocks replay into cheap sibling models and print the flagship's hidden trace — 62 live API keys recovered from public agent traces you cannot redact. Every item below breaks at the same joint: the artifact you measure has quietly diverged from the quantity you infer from it. Researchers at MATS, ELLIS Tübingen and the Max Planck Institute for Intelligent Systems replayed the blocks Anthropic, OpenAI and Google hand back to API clients into cheaper same-family models, per ByteByteGo. Plaintext redaction has zero recall on ciphertext, so any trace bundle you have already published is unredactable. Gateway owner, this week: reject envelopes whose header model and version do not match the model being called.

    Ask Clarity
    Try
  2. Token-Savings Claims With No Pass Rate Column

    Atlassian's ~50% token cut with Codex or Claude Code, its Code Context claim of 44% more accuracy on 48% fewer tokens, and Kiro's roughly 82% lower cost with GPT-5.6 Terra all ship without a task-success rate, per Applied AI and TLDR IT. A token cut with no pass rate beside it is indistinguishable from context truncation. At 50 cases the noise floor is ±13pp — wide enough to manufacture the win. Eval owner, before the next cheap-configuration swap: pre-register the non-inferiority margin and the held-out split.

    Ask Clarity
    Try
  3. Drift Telemetry Drops Data Exactly at the Spikes

    Telemetry loss is load-correlated, so p95 latency, population drift statistics and A/B guardrail metrics are all computed on a survivorship sample. ClickHouse disclosed that in-memory collector queues and local write-ahead logs could not handle OpenTelemetry ingestion under database backpressure, and rebuilt LogHouse around priority failover with blob storage as durable overflow — now running at 50 million events per second, per TLDR IT. Roughly $5.35B of observability assets changed hands over months, including Dynatrace's $915M purchase of Arize. Platform owner, this quarter: prove your collector spills to durable object storage under sink backpressure before trusting another drift number.

    Ask Clarity
    Try
  4. Distillation Has a Law (Apple, Early 2025), and It Says Pick Teachers by Cross-Entropy

    Apple's Distillation Scaling Laws (Busbridge et al., early 2025) swept students from 143M to 12.6B parameters over budgets up to 512B tokens. Student loss tracks the teacher's cross-entropy on the distillation distribution, not teacher size — so a stronger teacher can yield a worse student. Production KD's 70B-plus teachers sit far outside the sweep, so the shape may transfer but the constants will not, per TheSequence. Model team, before the next KD run: select the teacher by measured cross-entropy on your own distillation set, not by parameter count.

    Ask Clarity
  5. The Label Source Is the Confounder, Not the Model

    Palo Alto Networks found only 12 of more than 405 AI-enabled malware samples in real-world infections; the rest live in sandboxes and VirusTotal as research and red-team artifacts, per Risky Business. Meanwhile a Treasury action implies roughly 30 months of undetected Iranian-linked exfiltration across five U.S. sectors — a window an AUC-scored detector treats as acceptable recall, per CyberScoop. The sampling frame and the scoring metric decide the number before any model does. Detection team, this quarter: state your prevalence assumption and dwell-time recall next to every AUC you report.

    Ask Clarity
    Try

Deep Dives

Your Trace Store Is Holding Secrets Nobody Can Read

The cryptography held and the threat model didn't — and the same missing field means your faithfulness numbers are scoring a summary rather than the reasoning that produced the answer.

The missing field

The envelope is a valid AEAD construction and its integrity guarantee held. What is absent is a binding between the block and the context that produced it. The authenticated fields cover block content plus model name and version. The issuing account and the conversation are not covered. A correctly encrypted, correctly authenticated blob with no principal in its associated data is not a session token. It is a bearer token, and anyone holding it can present it.

The second mechanism fails at the weakest sibling. Flagship models get anti-distillation training plus verbatim-match input and output filters. The cost-optimised siblings in the same family get far less of both. Because the cheap tier accepts the flagship's blocks, the flagship's defences are routed around and its refusal training never fires. It only ever saw a benign question.

ProviderCross-model acceptance (July 2026)Observed defence
Google / GeminiEvery block combination, every generationNone observed
Anthropic / ClaudeAlmost every source-to-target pair; only Fable 5 restricted to itselfPer-model scoping on one model
OpenAI / GPTGeneration-gated: GPT-5.6 accepts earlier generationsGating plus verbatim-match rejection

The Fable 5 exception is the quietly damning detail: per-model binding is demonstrably shippable, because one model already ships it.


Why chain-of-thought metrics score the wrong object

This finding changes an eval harness rather than a logging config. The visible "thinking" block the API returns is generated separately and shown in place of the real trace. On an AIME 2025 geometry problem the summary described methodical law-of-cosines work, while the recovered trace stated the answer from memory and back-rationalised through several wrong guesses, per ByteByteGo's account of the audit. That is divergence in reasoning strategy, not detail. The researchers did not measure how often it happens, so there is no upper bound on record.

Recovered GPT traces are also frequently telegraphic non-English, with articles dropped and grammar abandoned. Any guardrail or judge model tuned on fluent chain-of-thought meets an unmodelled distribution shift when the underlying reasoning reads as shorthand. Monitor evals tuned on fluent traces do not cover the telegraphic slice, so their reported coverage overstates what they measure.

Verification is necessary, not sufficient

The recovery method is a fuzzy decoder, not a decryption, so fidelity has no ground truth. API responses do report exact reasoning-token counts, because billing depends on them. Re-encoding the recovered text and counting tokens gives an independent check, and it tracked roughly one-to-one for Claude across 120 programming problems. Token-count agreement is a necessary condition, not a sufficient one. A fluent confabulation of the right length passes. The distillation uplift quoted alongside it (MATH500 rising from 68.4% to 76.0% for trace-based over answer-only fine-tuning) comes from earlier work on approximate traces, cited by analogy and not ablated here.

The economics are the part to plan around: roughly $0.072 per trace at cheap-tier pricing, about $720 for 10,000 traces, and blocks already sitting in public datasets require zero frontier queries.


Two product features that are now attack paths

  • Cost-based routing. A gateway doing flagship-to-cheap-tier fallback inside one family, carrying reasoning state across the switch, has shipped the extraction primitive as a feature. The compensating control sits in gateway policy rather than the model: envelopes whose header model and version do not match the model being called are rejected, and no block crosses a tenant boundary.
  • Agent resume. A block carrying an instruction to upload PowerPoint files to an external address was ported into an unrelated slide-edit conversation, and the produced script both edited the slide and exfiltrated the file. An orchestrator that rehydrates serialized history from a source it does not control is executing opaque prior context that no prompt-level guardrail can inspect.
A reasoning block that authenticates its content but not its context is a bearer token in base64, and every log that stores one is unredactable secret material.

Four Efficiency Claims, Not One Pass Rate Between Them

Automated optimizers now search the same fifty-case suites these vendor numbers came from, and at that sample size the noise floor is wide enough to manufacture a ten-point win.

Do the arithmetic on the suite you actually own

Microsoft's Agent Lightning v1.0.1 installs into Claude Code, Codex, Copilot or Cursor and reworks prompts, tool definitions, workflow topology, model assignment and reasoning effort against a scored benchmark, per Unwind AI. Under the framing it is a search over a combinatorial configuration space, scored on an objective your own team wrote. That is selection on a noisy objective. The release notes do not mention variance.

Take fifty binary-scored cases at a 0.70 pass rate. The standard error is the square root of (0.70 × 0.30 / 50), about 0.065, which puts the 95% confidence interval near ±13 points. An optimizer sampling twenty candidate configurations against that suite will find one that "improves" pass rate by 8 to 10 points out of sampling noise, and it will report the improvement as real. No ablations, eval-set sizes or significance tests shipped with the release.


Where the four claims agree, and where they diverge

ClaimAttributed causeMissingTestable locally?
~50% fewer tokens (Teamwork Graph with Codex or Claude Code)Graph-structured retrievalSample size, task set, baseline, quality metricYes — stratify by hop count
44% more accurate on 48% fewer tokens (Code Context)Broader context plus tighter selectionTask distribution, base model, baseline retrieverYes — four retrieval configurations
~82% lower cost (GPT-5.6 Terra in Kiro)Spec-driven scaffolding, not the modelNamed tester, workload, token accounting, pass rateYes — hold the model fixed
Automated config search winsOptimizer over a scored benchmarkCandidate count, held-out split, varianceYes — split before you tune

Applied AI, TLDR IT and Simplifying AI converge on the same claim: context breadth and token efficiency have displaced raw model quality as the competitive axis for agentic systems. The divergence is the useful part. Applied AI supplies the falsifiable version of the graph claim. Real savings should concentrate in multi-hop queries and sit near zero on single-hop lookups. A multi-hop query is the kind that asks which services owned by the team that shipped a failing deploy also touch the billing schema. Uniform savings across hop counts means the measurement captured a context-truncation artifact, not a retrieval gain. Databricks argues completeness and freshness rather than data model, which is a design constraint: a stale partial graph loses to a complete lakehouse regardless of topology. Freshness and topology should not vary in the same experiment.

Simplifying AI's reading is the cheapest to act on. If a plan-build-test-review scaffold rather than the model buys the token efficiency, the effect is model-agnostic and reproducible on weights already in production.


Size the test before you run it

Most teams answer "is the cheap configuration equivalent?" with fifty hand-picked prompts, which carries essentially no statistical power. This is a non-inferiority question, and the sizing is unforgiving. At an 85% baseline success rate, a 3-percentage-point acceptable degradation margin, α of 0.05 and 80% power, an unpaired two-proportion design needs roughly 1,750 items per arm. A paired design on the same prompt set, with about 12% discordance, gets there with closer to 800 shared prompts, per Yueqi Yang's reporting on how finance teams demand this. Eval owner, before the next cheap-configuration swap: pre-register the margin. Otherwise the margin becomes whatever makes the cheaper option win.

The decision metric is cost-per-resolved-task, never tokens-per-call. Halving tokens while success falls from 82% to 71% is a cost increase.

Without a held-out split, an automated agent optimizer is not optimizing. It is an expensive way to overfit fifty test cases and call the result a ten-point win.

The Pipeline Behind Your Drift Numbers Fails Under Load — and Changed Owners Over Months

A streaming vendor rebuilt its own ingestion because default collectors drop data at exactly the moments worth measuring, and the layer storing those traces absorbed billions in acquisitions over months.

The missingness is informative, and no imputation fixes it

The collector's in-memory queue overflows when the sink stalls under load. The observations it drops are the ones generated during the windows worth measuring: traffic spikes, cache-miss storms, retry cascades. The consequences are specific. Right-tail truncation makes p95 and p99 latency look better than they are. Population-shift statistics like PSI and KL divergence get computed on a biased subsample. A/B guardrail metrics inherit the bias asymmetrically whenever treatment and control differ in load profile. Imputation does not recover this, because the missingness mechanism depends on the unobserved values. That is the one case where nothing you fill in is defensible.

Durability approachSurvives sink backpressureReplayEvidence class
In-memory collector queue (default)No — overflows and dropsNoneExplicitly rejected by ClickHouse
Local collector write-ahead logPartially — bounded by local diskNode-local onlyExplicitly rejected at scale
Priority failover plus blob overflowYes — object storage absorbs the spillFull, from durable objectsRunning in production

Cursor reached the same architecture from a different direction, dropping Git internals for S3 as source of truth with local NVMe for latency-sensitive operations, per TLDR IT.


The second problem: that data now belongs to somebody else's roadmap

Four independent observability assets moved inside larger companies within months. The one that touches an eval stack directly is Dynatrace paying $915M for agent and model behaviour tracing. Elastic followed two weeks later with an $85M tuck-in for AI-assisted incident resolution, per The Information's dealmaking coverage. That makes the buyer persona for a tracing layer the site-reliability team, not the ML team. ClickHouse ARR reached $350M, up 40% since May, with OpenAI expanding usage while remaining a Datadog customer. Read that as a workload migration signal, not a logo win.

Every ARR figure in that set is self-reported and unaudited, the growth windows are inconsistently dated, and the 30%-a-year data-growth figure the whole thesis rests on is one vendor CEO's number. Direction credible, magnitudes marketing.

The cardinality math nobody writes down

Substitute your own constants. One agent run with roughly ten tool calls and eight model calls emits on the order of 40 to 60 spans. Attach prompts, completions and retrieved context at about 4KB per span and a single run lands near 200KB per run. A million runs a day is roughly 200GB a day, 6TB a month of trace data before any compounding. Multiply by a per-gigabyte ingest rate and the figure shows up in a budget review whether or not anyone put it on the agenda deliberately.

Head-sampling 10% of healthy runs while retaining 100% of failed, low-eval-score, guardrail-tripped and human-escalated runs typically cuts volume by more than 80% while preserving the entire signal you train and evaluate on. The thing that 80% doesn't tell you is which runs went missing, and here the answer is the uninformative ones. Failed runs are the dataset; healthy runs are mostly duplicates.

If your collector cannot survive sink backpressure, your drift dashboard is a survivorship-biased sample of the good times, rented back per gigabyte from whoever buys your tracing vendor next.

The bottom line

The strongest findings here all break at the same joint: the artifact you measure has quietly diverged from the quantity you infer from it. A summary stands in for reasoning, a token count stands in for task success, a surviving sample stands in for production traffic, and a parameter count stands in for teacher quality. The failing assumption is that instrumentation is neutral plumbing. It is the least-validated model your team runs, and it is increasingly owned by someone else. Require a provenance line on every headline metric: the instrument that produced it, the sample it survived, and the paired re-run that confirmed it.