The 56% Replicates. The 2% Doesn't.
AT&T's cascade is the first enterprise cost/quality frontier anyone can copy, but the savings arrive as a mean while the risk concentrates in the hardest decile of your traffic.
Start with the denominator
Forty percent of queries is not 40% of tokens, and it is a long way from 40% of dollars. Summarization prompts are short. Agentic code generation is long-context and output-heavy. The disclosure carries its own tell: developers still send generation work to frontier models and use the cheap open models to summarize previously submitted code, the highest-volume and lowest-value slice of the traffic. Query share can rise forty points while spend barely moves.
The cost baseline is softer than the headline too. The Information's "at least hundreds of millions annually" is extrapolated from public list prices. Work it back: 45B tokens/day is roughly 16.4 trillion tokens a year, which implies a blended rate above $12 per million tokens. That is frontier output pricing with no enterprise discount and no prompt caching. A business case anchored to that multiple will not survive the first invoice.
The router is a production classifier, and nobody evaluated it
Routing "based on the complexity of the task" means a learned or heuristic classifier is sitting on the hot path with no reported eval. Treat it as any other production model: a labeled routing set, a confusion matrix with explicitly asymmetric misclassification costs, a confidence threshold that escalates instead of guessing, and drift monitoring on the input distribution. Onboard a new team or a new repo and the difficulty distribution shifts, so the decision boundary degrades quietly. The only symptom is a slow rise in retries.
The levers that compound underneath it
| Lever | Reported effect | What to verify locally |
|---|---|---|
| Complexity routing (LiteLLM) | -56% coding cost, -2% mean quality | p95 regression per task class, not the mean |
| Semantic cache | 57.1% hit rate, 55.7% fewer tokens, ~15ms hit latency | False-hit precision — the similarity threshold is a correctness decision |
| Gisting / context compression | ~40% lower end-to-end latency, ~15% higher throughput (per Shopify's writeup) | Quality at your compression ratio |
| CacheBlend, now open source in LMCache | 2-4x on multi-document queries, order-independent | p50/p95, tokens recomputed, answer quality on your eval set |
One of these is an active bug rather than an opportunity. Prompt caching on OpenAI and Anthropic matches on an exact byte-for-byte prefix. Reorder two cached documents and both become misses. Cache document A and document B separately, then query them together, and B misses because its key-value state was computed without ever seeing A. In production the per-block hit distribution is bimodal: a small fraction of blocks serves almost all the hits while the rest are dead weight, so any saving reported off a mean hit rate needs re-deriving. A reranker exists precisely to vary chunk order per query, which is structurally incompatible with prefix caching. CacheBlend recomputes only the cross-boundary tokens and reuses the rest, the first design here that reconciles the two. The thing its 2-4x doesn't tell you is the hardware, the document-length distribution, or the quality metric, since none of the three is named. The mechanism justifies a spike. The number is yours to reproduce.
Change the unit before you change the tier
ARC Prize is now quoting cost per task rather than cost per token. Gemini 3.7 Flash lands 84.6% on ARC-AGI-2 at $0.25 per task. For agentic work that is the only honest denominator, because a cheap model that burns three times the rollouts is not cheap. Rebuilt on that axis, the Pareto curve may name a different default tier than the one a token-price dashboard picked.
A 56% cost cut for a 2% quality drop is only a good trade if you can prove the 2% is not concentrated in your hardest 5% of tasks.
What to do
Run a two-to-four week shadow-mode dual-inference study before shifting any production traffic: replay a stratified sample of prompt logs against the incumbent frontier model and 2-3 open-weight candidates, with a pre-registered non-inferiority margin of 3% relative pass rate per task family.
Instrument per-block cache hit rate in the LLM gateway this sprint, plot the distribution rather than the mean, and enforce a canonical ordering of cacheable blocks at the front of every prompt in code.
Replace the '% of queries on open models' dashboard with token-weighted cost per accepted task and p95 quality regression per task class by the end of the quarter.