Your Trace Store Is Holding Secrets Nobody Can Read
The cryptography held and the threat model didn't — and the same missing field means your faithfulness numbers are scoring a summary rather than the reasoning that produced the answer.
The missing field
The envelope is a valid AEAD construction and its integrity guarantee held. What is absent is a binding between the block and the context that produced it. The authenticated fields cover block content plus model name and version. The issuing account and the conversation are not covered. A correctly encrypted, correctly authenticated blob with no principal in its associated data is not a session token. It is a bearer token, and anyone holding it can present it.
The second mechanism fails at the weakest sibling. Flagship models get anti-distillation training plus verbatim-match input and output filters. The cost-optimised siblings in the same family get far less of both. Because the cheap tier accepts the flagship's blocks, the flagship's defences are routed around and its refusal training never fires. It only ever saw a benign question.
| Provider | Cross-model acceptance (July 2026) | Observed defence |
|---|---|---|
| Google / Gemini | Every block combination, every generation | None observed |
| Anthropic / Claude | Almost every source-to-target pair; only Fable 5 restricted to itself | Per-model scoping on one model |
| OpenAI / GPT | Generation-gated: GPT-5.6 accepts earlier generations | Gating plus verbatim-match rejection |
The Fable 5 exception is the quietly damning detail: per-model binding is demonstrably shippable, because one model already ships it.
Why chain-of-thought metrics score the wrong object
This finding changes an eval harness rather than a logging config. The visible "thinking" block the API returns is generated separately and shown in place of the real trace. On an AIME 2025 geometry problem the summary described methodical law-of-cosines work, while the recovered trace stated the answer from memory and back-rationalised through several wrong guesses, per ByteByteGo's account of the audit. That is divergence in reasoning strategy, not detail. The researchers did not measure how often it happens, so there is no upper bound on record.
Recovered GPT traces are also frequently telegraphic non-English, with articles dropped and grammar abandoned. Any guardrail or judge model tuned on fluent chain-of-thought meets an unmodelled distribution shift when the underlying reasoning reads as shorthand. Monitor evals tuned on fluent traces do not cover the telegraphic slice, so their reported coverage overstates what they measure.
Verification is necessary, not sufficient
The recovery method is a fuzzy decoder, not a decryption, so fidelity has no ground truth. API responses do report exact reasoning-token counts, because billing depends on them. Re-encoding the recovered text and counting tokens gives an independent check, and it tracked roughly one-to-one for Claude across 120 programming problems. Token-count agreement is a necessary condition, not a sufficient one. A fluent confabulation of the right length passes. The distillation uplift quoted alongside it (MATH500 rising from 68.4% to 76.0% for trace-based over answer-only fine-tuning) comes from earlier work on approximate traces, cited by analogy and not ablated here.
The economics are the part to plan around: roughly $0.072 per trace at cheap-tier pricing, about $720 for 10,000 traces, and blocks already sitting in public datasets require zero frontier queries.
Two product features that are now attack paths
- Cost-based routing. A gateway doing flagship-to-cheap-tier fallback inside one family, carrying reasoning state across the switch, has shipped the extraction primitive as a feature. The compensating control sits in gateway policy rather than the model: envelopes whose header model and version do not match the model being called are rejected, and no block crosses a tenant boundary.
- Agent resume. A block carrying an instruction to upload PowerPoint files to an external address was ported into an unrelated slide-edit conversation, and the produced script both edited the slide and exfiltrated the file. An orchestrator that rehydrates serialized history from a source it does not control is executing opaque prior context that no prompt-level guardrail can inspect.
A reasoning block that authenticates its content but not its context is a bearer token in base64, and every log that stores one is unredactable secret material.