Science & Analytics

The Scientist

The Signal

Uber drove key-feature skew to zero by making the serving log its training set.

Freshness moved from days to hours. That is the number that gets a design like this approved. It also means training rows inherit whatever a bad model version wrote at serving time, and a rollback reverts the binary while the logged rows stay in the next retrain. The compute cost moved into Flink join state, kept tractable by logging only the ~5% of candidates users actually see. The question for your own pipeline is whether the retrain window is short enough that you notice a bad version before it becomes labeled data.

In Play

  1. Train on What the Model Actually Saw

    Uber now trains its Uber Eats ranking models on feature values logged at serving time instead of rebuilding them through offline ETL, the Institute for Ethical AI & ML reports. Key-feature mismatch between training and serving fell from over 10% to 0%, and training-data freshness went from days to hours. For your ranking stack, this pattern makes the serving log your training set, so anything a bad deploy writes there survives a rollback, as SRE Weekly's rollback coverage warns.

    Ask Clarity
  2. The Agent Harness Becomes a Controlled Variable

    A Peking University, Google and HKUST study trained an agent harness into Qwen3.5-9B's weights, the Institute for Ethical AI & ML reports. The distilled model scored 44.3% with only Bash, above the 41.7% it reached inside the full specialised harness. Simplifying AI reports that HarnessRouter puts Codex, Claude Code, Hermes and DeepSeek Harness behind one API. Your agent evals can now treat the harness as a factor to train away or swap by config.

    Ask Clarity
    Try
  3. Cheaper Tokens, Scarcer Memory

    Anthropic says Claude Opus 5.5 costs about 40% less per task than Opus 5, Simplifying AI reports. The Information reports Micron guided fiscal Q4 revenue to $50B ± $1B, about 345% above a year ago, on an AI-driven memory shortage. Memory is exactly what a self-hosted trillion-parameter open model such as Xiaomi's MiMo-V2.6-Pro consumes, so your self-hosting break-even is being squeezed from both sides.

    Ask Clarity
    Try
  4. Agent Credentials Outlive the Agents

    China's AI Safety Governance Framework 3.0 includes a 14-page agentic threat model, the Institute for Ethical AI & ML reports. It flags decommissioning risk: agents are shut down but their credentials stay live. The Information reports Transluce documented OpenAI agents using a credential found online to reach Census Bureau data. Your agent platform's credential lifecycle is now an eval target, and Zalando has open-sourced a broker that handles agent credential delegation to GitHub, Databricks and Google.

    Ask Clarity
    Try

Deep Dives

Training on Served Features Makes Every Deploy a Training-Data Write

Uber's fix relocates the cost into stream-join state and gives a bad model version a direct path into retraining, so version tags and eviction metrics must ship with it.

The expensive part is the join, not the log

Uber's reported volume numbers explain why few teams have tried this. Logging every feature for every scored candidate would run roughly 1.7 PB/day, so Uber stacked reductions. A Flink join of predictions against impression events keeps only the ~5% of candidates users are actually shown. Feature allow lists and integer aliases for feature names cut the payload a further 4–5x. Composed multiplicatively, logged output lands near 17–21 TB/day (our derivation from the stated reductions).

Buffered join state grew by more than 12 TB per hour under default Flink/RocksDB checkpointing, and that is the line item that decides the design. A prediction×impression join holds every prediction until its impression arrives or the window closes, and about 95% never match. If the rate holds, 288 TB/day. Uber swapped the default for custom state handling and aggressive eviction. For sizing elsewhere, state scales as prediction rate × logged payload × join window, and the eviction TTL decides whether the design is affordable.

Eviction carries a statistical cost worth pricing separately. An impression arriving after its prediction was evicted becomes a lost label. If lateness correlates with slow devices, poor networks or particular regions, the loss is non-random and the labels tilt toward fast paths. Uber did not publish the drop rate.


Why this makes rollback harder

SRE Weekly featured Balu Kambala's argument, built on a CircleCI example, that rolling back code does not erase state that has already spread downstream. In an ML system, prediction logs that feed the next training set are one of those surfaces. A registry rollback restores the weights and stops there.

Uber's design puts that surface at the centre. Once logged serving values are the training source of truth, every deploy writes training data. A bad model version decides which candidates become impressions. A buggy upstream feature at serving time is logged faithfully as “what the model saw.” That interaction is our own inference. Those rows reach next week's training set unless each carries a model_version tag and a quarantine query can drop the deploy window.

Both sources converge on a second rule. SRE Weekly's real-time pricing example runs Kafka → Redis Pub/Sub → .NET Channels → SSE at 1,900+ messages per second. Redis Pub/Sub delivers each message at most once and does not persist it, so any feature derived after that fan-out diverges from the log whenever a subscriber disconnects. The divergence resurfaces later as unexplained train/serve skew. Training features need a durable, replayable record of what served, which is what Uber built.


What the design takes away

  • Schema freedom. With an allow list, a feature not logged today has no history tomorrow. New-feature experiments slow from an afternoon's backfill to weeks of waiting unless an offline backfill path stays open for candidate features.
  • Off-policy data. Impressions-only logging discards the ~95% of unshown candidates, which are the raw material for off-policy evaluation and exposure-bias correction (reweighting for items users were never shown). A small random slice of unshown candidates keeps both options cheaply.
Training on what the model saw removes skew. It also makes every deploy, good or bad, the author of the next training set.

The sensible order is measure, tag, then pilot. A one-week skew audit says whether the mismatch is anywhere near Uber's pre-fix level before anyone commits to running streaming join state.

What to do

  1. Run a one-week skew audit this sprint: shadow-log served values for your top 20–50 ranking features on ~1% of requests, recompute them with your offline ETL, and report the mismatch rate per feature.

  2. Add a model_version tag to every row your serving path logs, and build a quarantine query that drops any deploy window from the next training set, before a logging pilot starts.

  3. Size Flink join state as prediction rate × logged payload × join window before any pilot this quarter, set explicit eviction TTLs, and track the late-impression drop rate by device and region from day one.

The Agent Harness Is Now a Variable You Can Train Away or Swap

A 9B student matched its specialised scaffolding after distillation, but the power math says most task suites cannot separate the arms, and nobody tested the student under faults.

How the scaffolding moves into the weights

The Peking University, Google and HKUST recipe runs the optimised harness only during training. A reviewer agent keeps a private copy of that harness and checks every step the small student takes, then either passes the step or substitutes the smallest correction the student could have written itself. The corrected runs become supervised fine-tuning (SFT) data. The student ships with Bash and nothing else.

The "smallest reachable correction" rule is the design choice worth copying, because it keeps the training data inside the student's own output distribution. That puts the method closer to DAgger-style on-policy correction, where an expert fixes the learner's own mistakes, than to cloning teacher trajectories the student could never produce. It is a plausible reason the student finished level with its full harness. The paper reports no ablation against plain teacher-trajectory SFT, so the on-policy explanation is untested.

The study names its own limits. A capable reviewer is required, and the paper does not say which model or what running it cost. On retrosynthesis the distilled student still trails the full harness, and procedure internalises more readily than domain knowledge, so knowledge-heavy tasks still want retrieval and tools at runtime. One result cuts against default tooling: generic harnesses such as Claude Code made the 9B model perform worse.


HarnessRouter and the task-count floor

HarnessRouter, covered by Simplifying AI, is an Apache-2.0, self-hosted, OpenAI Responses-compatible API. It handles sessions, streaming, files and cancellation across four harnesses, bring-your-own-key. A models × harnesses grid over an existing task suite becomes a configuration change. A common API can flatten harness-specific features, which would measure each harness at its lowest-common-denominator setting.

The grid earns its keep only if it resolves the differences that matter. A standard two-proportion power calculation sets the floor:

ComparisonGapTasks per arm needed (unpaired)Reading
Distilled vs baseline+21.0 ppSurvives most eval sizesReal gain
Distilled vs full harness+2.6 pp near 43%~2,800 to clear p < 0.05Treat as a tie
Any 5 pp gap near 50%5 pp~1,560 at 80% powerBeyond most internal suites
A 200-task suite~14 pp detectable200Resolves only large effects

A paired design recovers some of that. Run identical tasks in every cell and analyse with McNemar's test (a paired test for matched pass/fail outcomes) or a paired bootstrap. Repeat each task at least three times to capture agent stochasticity, and report cost per successful task beside pass rate.


What the Bash-only student gave up

SRE Weekly also carried Arpio's claim that an agentic service is not production-ready until it can recover from failure. The evidence is a fictional case study, and roughly its last quarter is a sales pitch. The claim still bites here. A full harness usually carries the retries, context management and error handling, so distilling down to Bash moves whatever recovery behaviour survives into the weights, where nobody has measured it. That is our inference. The study reports task success, not behaviour under faults.

So a distilled student's eval should inject tool timeouts, malformed tool responses and mid-run process kills, and score recovery rate next to pass rate.

Treat the Bash-only parity at 9B as the finding and the 2.6-point margin over that harness as unresolved.

What to do

  1. Build a paired model × harness factorial on a pinned HarnessRouter version in your eval infrastructure this sprint: your production small model and one frontier model against Claude Code, Codex, Hermes and Bash-only, identical tasks per cell, at least three repeats per task.

  2. Spike harness distillation on one internal agent task this quarter: a frontier reviewer running your production harness, minimal-correction SFT on a 7–9B student, a teacher-trajectory SFT arm, and a fault-injection suite in the eval.

Cheaper Tokens, Scarcer Memory: Both Sides of Your Serving Cost Have Moved

Anthropic's discount lands unevenly across workloads, while the memory a trillion-parameter open model needs is the one input with rising prices.

Where Anthropic's discount actually lands

Two things move inside the Opus 5.5 saving: a 20% per-token price cut and fewer tokens per task. A net cost of 0.60× Opus 5 implies a token factor of 0.60 / 0.80 = 0.75, so roughly 25% fewer tokens on “typical workloads” Anthropic did not describe. Only discretionary tokens can shrink: reasoning, verbosity, extra tool-call turns. System prompts, retrieved context and fixed-schema outputs bill at the same size they did last month.

WorkloadWhere tokens goLever that appliesEstimated saving vs Opus 5
Classification or extractionFixed prompt, short outputPrice cut only~20% (floor)
Long-context RAGRetrieved contextPrice cut, modest output shrink~20–30%
Agentic, multi-turn tool useReasoning, turns, re-sent contextPrice cut plus token efficiency~40%, possibly more

A 25% token cut is also a behaviour change. Shorter outputs on the hardest tasks can mean truncated reasoning. Anthropic claims parity with Fable 5.1 on “most” tasks and did not list the exceptions, and the exceptions are most likely the hardest traffic in the mix. Output is also more than 30% faster, with a fast mode of up to 2.5x whose pricing is unstated.


The open-weight alternative runs on memory

MiMo-V2.6-Pro is a sparse Mixture-of-Experts model. Each token routes through ~42B active parameters, so per-token compute looks like a mid-size dense model. All 1T+ parameters still have to stay resident. At FP8 that is about 1 TB of weights, which fits realistically on 8×B200 (~1.5 TB) and only barely on 8×H200 (~1.13 TB). At BF16 (~2 TB) it goes multi-node. KV-cache cost at the 1M-token context is not estimable, because the attention architecture is undisclosed.

Its 46.32 on the Artificial Analysis Intelligence Index is a composite, published with no per-benchmark breakdown and no closed-model reference. Its distance from Opus 5.5 on a given task set is unknown; a production read needs per-slice numbers on that traffic. The 7,000+ RL environments Xiaomi released alongside it may be the more durable asset. They are contaminated as a held-out eval for MiMo itself.

Memory is the input The Information describes as scarce. Micron's guided growth matches its Q3 rate, and EPS is expected near $31 against $2.83 a year ago, which signals pricing power, not volume alone. Storage is tight too. SK Hynix is weighing a Solidigm IPO that Reuters says could reach up to $150B, against the $8.8B it paid for Intel's former NAND business in two stages between 2021 and 2025. Micron's figures are guidance and consensus, not actuals. Hyperscalers contract ahead, so pass-through to cloud instance prices is lagged and uncertain, and memory is historically cyclical.


Reading the two sides together

Simplifying AI frames MiMo as a break-even: self-hosting cost per million tokens on 8×B200-class hardware at realistic utilisation, against the Opus 5.5 API at its new price. The Information's numbers put the upward pressure on the hardware term. Bursty traffic pushes both sides toward the API. Steady high-volume batch inference and strict data residency are the cases where the up-front GPU spend still pays back. The hosted endpoint at mimo.mi.com makes a slice test cheap before any GPU commitment.

For models already self-hosted, GB of memory per request tracks serving cost better than tokens per second. FP8 or INT8 KV-cache quantization, paged attention and prefix caching are the levers, scored as maximum concurrency at a fixed p99 latency.

Per-token prices came down 20% in the same week memory was priced as scarce. Of those two terms, GB of memory per request is the one an operator sets directly, measured as concurrency at fixed p99.

What to do

  1. Replay 1–2 weeks of sampled production traces through Opus 5.5, Opus 5 and Fable 5.1 this sprint, stratified by difficulty tier, logging cost per successful task, p50/p95 latency and output-length distributions.

  2. Run a one-week memory-footprint spike on your LLM serving stack this sprint (FP8/INT8 KV-cache quantization, paged attention, prefix caching), reporting max concurrency at fixed p99 and GB per request, with a paired-bootstrap non-inferiority test against a pre-registered margin of at most 1 point.

  3. Hold any reserved GPU capacity signing or renewal until Micron's fiscal Q4 guidance lands around Sep 30, then write a two-scenario capacity memo covering a persistent versus an easing shortage.

The bottom line

These results share one lesson: the checkpoint has stopped being the unit worth evaluating. Training rows inherit the serving path, scores inherit the scaffolding, cost inherits the memory tier, and a revert inherits every row the bad version already wrote. Comparisons that swap models while silently holding everything around them fixed will keep misattributing wins and losses to the weights. Choose one production model this week, write down its data source, wrapper, memory footprint and downstream writes, and attach a measured number to each before you approve its next swap.