Science & Analytics

The Scientist

The Signal

AT&T bought a 56% coding cost cut with a 2% quality drop it never measured.

Forty percent of queries on open models is not forty percent of tokens, and it is further still from forty percent of dollars. The cheap tier summarizes code the frontier models are still writing. An unverified mean can also hide a steep regression concentrated in the hardest slice of traffic, which is where rework hours actually accrue. If the routing mix in your own cost model looks like this one, the savings and the quality figure are being measured on different populations.

In Play

  1. Production Routing Finally Has a Reference Number

    AT&T's VP of AI, Mark Austin, told The Information that a LiteLLM complexity router cut AI coding costs 56% for a stated 2% quality decline. Open-weight models now serve 40% of employee queries at 45 billion tokens per day. Deep dive below on why the cost cut transfers and the quality figure doesn't.

    Ask Clarity
    Try
  2. The Aggregate Passed, the Missing Arm Failed

    All 33 blood-glucose prediction models in one review looked equivalent in aggregate, yet every one ran roughly 6 mg/dL worse for Type 1 than Type 2 patients. Separately, an SSRN working paper reports that generative AI raised pupils' homework scores 18% while their closed-book exam scores fell 20%. Deep dive below on the shared defect and the release gate that catches it.

    Ask Clarity
    Try
  3. MLflow Joins the Actively Exploited Column

    CISA confirmed active in-the-wild SSRF exploitation of MLflow servers. DataDog separately analysed N4D Mesh Controller, a Linux strain scanning for internet-exposed MCP servers. The same security reporting timed GitLab CVE-2026-19478 to exploitation 48 hours after proof-of-concept code dropped. A monthly patch cadence on internet-adjacent ML control-plane hosts is not defensible against that window. The exploitation chain and the inventory it implies are in the deep dive.

    Ask Clarity
    Try
  4. Semi-Structured Columns Go Native in Two Engines at Once

    DuckDB 2.0 and Iceberg v3 both shipped native VARIANT types on Parquet binary encoding, queryable from Spark SQL via variant_get, with timestamps, decimals, and binary values preserved inside the column. That retires the JSON-string-plus-cast-on-read pattern in model-input logging, where a silent cast failure surfaces three weeks into a training run. Netflix separately published a greater-than-99.99% dual-run agreement bar before retiring a legacy pipeline, which transfers verbatim as a feature-pipeline migration gate.

    Ask Clarity
    Try
  5. Compute Scarcity Is Priced in Time, Not Kilowatt-Hours

    ERCOT's June 2026 interruptible-load rule compresses interconnection from 5-7 years to 12-18 months for battery-buffered data centers that accept curtailment, and FERC has written to six other grids urging adoption. On a 1GW build, five years of energy is 4-9% of total cost of ownership while accelerators are roughly 70% of capex. That inverts the standard efficiency roadmap: model FLOPs utilisation and preemption tolerance are the cost levers, and joules per token is a carbon metric.

    Ask Clarity
    Try

Deep Dives

The 56% Replicates. The 2% Doesn't.

AT&T's cascade is the first enterprise cost/quality frontier anyone can copy, but the savings arrive as a mean while the risk concentrates in the hardest decile of your traffic.

Start with the denominator

Forty percent of queries is not 40% of tokens, and it is a long way from 40% of dollars. Summarization prompts are short. Agentic code generation is long-context and output-heavy. The disclosure carries its own tell: developers still send generation work to frontier models and use the cheap open models to summarize previously submitted code, the highest-volume and lowest-value slice of the traffic. Query share can rise forty points while spend barely moves.

The cost baseline is softer than the headline too. The Information's "at least hundreds of millions annually" is extrapolated from public list prices. Work it back: 45B tokens/day is roughly 16.4 trillion tokens a year, which implies a blended rate above $12 per million tokens. That is frontier output pricing with no enterprise discount and no prompt caching. A business case anchored to that multiple will not survive the first invoice.


The router is a production classifier, and nobody evaluated it

Routing "based on the complexity of the task" means a learned or heuristic classifier is sitting on the hot path with no reported eval. Treat it as any other production model: a labeled routing set, a confusion matrix with explicitly asymmetric misclassification costs, a confidence threshold that escalates instead of guessing, and drift monitoring on the input distribution. Onboard a new team or a new repo and the difficulty distribution shifts, so the decision boundary degrades quietly. The only symptom is a slow rise in retries.

The levers that compound underneath it

LeverReported effectWhat to verify locally
Complexity routing (LiteLLM)-56% coding cost, -2% mean qualityp95 regression per task class, not the mean
Semantic cache57.1% hit rate, 55.7% fewer tokens, ~15ms hit latencyFalse-hit precision — the similarity threshold is a correctness decision
Gisting / context compression~40% lower end-to-end latency, ~15% higher throughput (per Shopify's writeup)Quality at your compression ratio
CacheBlend, now open source in LMCache2-4x on multi-document queries, order-independentp50/p95, tokens recomputed, answer quality on your eval set

One of these is an active bug rather than an opportunity. Prompt caching on OpenAI and Anthropic matches on an exact byte-for-byte prefix. Reorder two cached documents and both become misses. Cache document A and document B separately, then query them together, and B misses because its key-value state was computed without ever seeing A. In production the per-block hit distribution is bimodal: a small fraction of blocks serves almost all the hits while the rest are dead weight, so any saving reported off a mean hit rate needs re-deriving. A reranker exists precisely to vary chunk order per query, which is structurally incompatible with prefix caching. CacheBlend recomputes only the cross-boundary tokens and reuses the rest, the first design here that reconciles the two. The thing its 2-4x doesn't tell you is the hardware, the document-length distribution, or the quality metric, since none of the three is named. The mechanism justifies a spike. The number is yours to reproduce.


Change the unit before you change the tier

ARC Prize is now quoting cost per task rather than cost per token. Gemini 3.7 Flash lands 84.6% on ARC-AGI-2 at $0.25 per task. For agentic work that is the only honest denominator, because a cheap model that burns three times the rollouts is not cheap. Rebuilt on that axis, the Pareto curve may name a different default tier than the one a token-price dashboard picked.

A 56% cost cut for a 2% quality drop is only a good trade if you can prove the 2% is not concentrated in your hardest 5% of tasks.

What to do

  1. Run a two-to-four week shadow-mode dual-inference study before shifting any production traffic: replay a stratified sample of prompt logs against the incumbent frontier model and 2-3 open-weight candidates, with a pre-registered non-inferiority margin of 3% relative pass rate per task family.

  2. Instrument per-block cache hit rate in the LLM gateway this sprint, plot the distribution rather than the mean, and enforce a canonical ordering of cacheable blocks at the front of every prompt in code.

  3. Replace the '% of queries on open models' dashboard with token-weighted cost per accepted task and p95 quality regression per task class by the end of the quarter.

Five Reported Wins, One Missing Arm

Subgroup gaps, tool-removed capability drops, self-grading judges, and argmax with free ground truth available are the same measurement defect wearing four different costumes.

The mechanism generalises past glucose

Type 1 physiology is more volatile, which means conditional variance differs by cohort. A loss function minimising average error will happily trade accuracy on the smaller, harder, higher-variance subgroup for accuracy on the easy majority, and a single-number scorecard records that trade as an improvement. Swap in new tenants against mature tenants, low-activity against high-activity users, or rare-event against common-event segments. The arithmetic does not change.

Reported resultWhat the aggregate hidThe arm that exposes it
33 glucose predictors, equivalent aggregate accuracy~6 mg/dL larger error for Type 1 in every model testedPer-slice error plus a max-minus-min gap threshold that fails CI
Generative AI raised homework scores 18%Closed-book exam scores fell 20% in the same populationUnassisted holdout plus a delayed tool-removed re-test
In-session LLM judge scores look strongThe evaluator inherits the generator's justification for its own errorsCleared-context judge, plus a deterministic oracle (execute the suite, don't read the diff)
Category classifier at 90%-plus top-1A photo printer emitted as "luggage"; one pool order split across Decor, Outdoors and GardenGroup-consistency metric and a calibrated abstention path
Cross-modal imputation for absent recordsOnly ~half the accuracy lost to a missing ECG or chest X-ray is recoveredModality/feature-group dropout ablation on your own models

The most credible number here is the unflattering one

Recovering roughly half of the accuracy lost to a missing modality is the strongest result in this set because it is disappointing. Papers reporting full recovery from a missing modality are usually leaking the signal back in through correlated features. Hold ~50% as the realism benchmark the next time a vendor argues that their imputation layer makes heterogeneous data availability a non-issue.

The productivity split needs the opposite treatment. The working paper reports no sample size, no standard errors, and no identification strategy, and it has not been peer reviewed, so take the design lesson and do not cite the point estimates. The design lesson is worth more than the estimates anyway: the assisted metric and the capability metric moved in opposite directions in the same population. Census self-report data rhymes at enterprise scale — 55% of US workers use AI, and 13% of users report no time savings or a net workload increase. Ticket resolution time, analysis throughput, and review latency are all measured with the assistant in the loop, and the delta gets booked as capability uplift. The thing those metrics do not tell you is what happens when the assistant is removed. The closed-book exam is the arm almost nobody builds, and the only one that separates leverage from dependency.


What changes in the harness

The engineering is small and the coverage gain is permanent. Enumerate the top five population slices, emit per-slice error alongside the aggregate, and add a max-minus-min gap threshold that blocks the release. Roughly a day of work, and it closes the failure mode that defeated all 33 glucose models.

Two adjacent gates cost about the same. Where a prediction overwrites a field already known with certainty — as in the order-confirmation case, where the verified item title sat in the order record — define a calibrated threshold below which the known value is emitted instead, then track abstention rate and post-abstention accuracy as first-class metrics. Where users see predictions in groups, add modal-label agreement or within-group label entropy to CI. Per-item accuracy and macro-F1 are both structurally blind to incoherent triples on the same basket.

A single-number scorecard cannot fail the subgroup it was averaged over.

What to do

  1. Add per-slice error plus a worst-group gap metric as a blocking CI gate on your top five population slices this sprint, and baseline it on 30 days of production traffic before touching the model.

  2. Re-score 200 already-judged generations with a cleared-context judge this sprint — no generator rationale, different exemplars, ideally a different model family — and report the paired delta with a bootstrap confidence interval plus agreement against human labels.

  3. Retrofit every AI-assistance eval you own with an unassisted holdout arm and a delayed tool-removed re-test by the end of the quarter, reporting both deltas side by side.

Your RAG Index Is Laundering the ACLs

A high-privilege crawler plus one shared vector index turns row-level access control into cosine similarity, and the failure is measurable long before it becomes an incident report.

The failure class DLP does not cover

In March 2026 an internal AI agent at Meta produced a Sev 1 by surfacing sensitive company and user data to employees with no entitlement to it. Not a jailbreak, not prompt injection. An authorization bypass that the access-control system scored as clean: a service account with legitimate broad access read data and handed it to a human.

That is the reference architecture of most internal retrieval deployments. One high-privilege credential crawls wiki, chat, tickets and the warehouse, one index holds all of it, nearest-neighbour results go to whoever asks. The index applied a lossy transformation to the permission model. Meta has a mature security organisation and a Sev 1 process, and it happened there anyway.

Make leakage a number, not a review item

Build a golden adversarial set of 600-plus (query, principal) pairs across privilege tiers, each with a labeled allow/deny document set, and define Unauthorized-Retrieval Rate as the fraction of returned chunks the requesting principal is not entitled to. Gate it at zero in CI on every index rebuild. Sample size is the whole argument: by the rule of three, zero failures in n trials bounds the true rate near 3/n, so 600 clean probes buys 0.5% at 95% confidence. A 50-query smoke test cannot separate "safe" from "leaks one in two hundred," and one in two hundred is the regime that produces an incident nobody predicted.

Then pay the retrieval bill knowingly. Filtering after approximate nearest-neighbour search spends top-k slots on documents that get discarded, so recall decays quietly as the entitlement pass rate drops, and any code path that skips the filter fails open. Filtered HNSW, per-tenant namespaces, or over-fetch at k' = k / expected pass rate. Then benchmark recall@10 per privilege tier before and after, so the accuracy cost of correctness lands in a design doc rather than a postmortem.


The control plane underneath it is now enumerated

CISA's confirmation of active SSRF exploitation against MLflow servers matters because of what a tracking host usually is: no auth, inside a VPC, holding an instance role that reads the artifact bucket. Instance role to artifact bucket to registry write to a poisoned model on every serving pod. Patch, add an authenticating proxy, enforce IMDSv2 with hop-limit 1, default-deny egress from tracking hosts. The same inventory question applies to every operator console that ships unauthenticated: Ray's job API on 8265, MLflow on 5000, tokenless Jupyter, Airflow, Label Studio, OpenAI-compatible vLLM or TGI endpoints. An open Ray job API is arbitrary code execution on your GPU fleet with a credential-rich environment attached. Hours of inventory, not a quarter of project.

The AI reviewer in front of it has unmeasured recall

Wiz's Red Agent exploited a vulnerable GitHub Actions workflow in a Snowflake repository that AI-based security checks had already cleared, and ended holding credentials to an unrelated SaaS. That is n=1 vendor research with no model named, no defect taxonomy, no recall data, so it establishes no rate. It does establish that the rate is unknown and non-zero, which is enough to reword the control in audit docs. Producing the missing number takes days: 50-100 labeled vulnerable workflow configs (over-permissioned tokens, unsafe triggers, unpinned third-party actions) through the reviewer, false-negative rate computed. Until that number exists, the model is a screen sitting under deterministic policy-as-code, not a gate.

A vector index is a permission-laundering layer until a number proves otherwise.

What to do

  1. Inventory every ML control-plane endpoint reachable beyond localhost, patch MLflow to current, place it behind an authenticating proxy, and enforce IMDSv2 with hop-limit 1 plus default-deny egress on tracking hosts.

  2. Add Unauthorized-Retrieval Rate to the RAG and agent eval harness with 600-plus labeled (query, principal) probes across privilege tiers, gated at zero in CI on every index rebuild, within the sprint.

  3. Measure your AI code reviewer's false-negative rate on 50-100 labeled vulnerable workflow configs before it gates another merge.

The bottom line

The pattern across today is narrower than "measure better": in every case the number that shipped was an average over a mixture the team could have split and didn't, and the split was cheap. That kills the convenient assumption that a vendor's effect size and your production risk sit on the same axis — they diverge exactly where your traffic is hardest, your cohorts smallest, and your evaluator most invested in its own answer. Pick the single number your next roadmap decision rests on and refuse to move until it ships with the slice, the holdout, or the clean-context re-score that could falsify it.