Science & Analytics

The Scientist

The Signal

Uber cut agent cost per request 34% and its total agent spend still rose 6.2x.

Requests grew 9.4x since February, which swamps a 52% drop in cost per session. A blended efficiency ratio won't tell you which of those two terms moved, and that is the number most teams report. Decompose it and each factor has a different owner: sessions to product, turns to the agent loop, tokens to prompts, price to procurement. When the bill moves next month, that decomposition is the difference between naming a lever and staring at an average.

In Play

  1. Unit Cost Fell 34%; Total Agent Spend Rose About 6x

    Uber's engineering team published a factor-level cost model for agentic workloads — users x sessions x turns x requests x tokens x price — reporting cost per 1,000 requests down 34% and cost per session down 52% from peak. Requests grew 9.4x since February, so absolute spend still rose roughly 6.2x. For you, that means every unit-economics win you report is compatible with a budget that tripled, and finance reads the second number.

  2. 770B Open Weights Land Inside Statistical Noise

    Tencent open-sourced Hy4 under Apache 2.0: 770B total parameters, roughly 49B active per token, 78 layers, context past 1M tokens, hosted at $0.834 per million input tokens. Its human-preference lead over Kimi K3 and GLM-5.3 is 0.05 points on a 4-point scale across 203 tasks judged by 163 experts, which is inside noise. Treat the three checkpoints as tied and let your own paired eval, not the release note, decide routing.

  3. Search Summaries Debunked Less Than Chat

    NPR posed 30 nation-state false narratives to six chatbots in both neutral and leading forms with web access enabled; the aggregate debunk rate was about 75%, and search-summary surfaces performed worse than conversational chat on the same claims. Same model class, different pipeline. Your product surface is the retrieve-and-summarize one, so a factuality number measured at a bare model endpoint describes a system you do not ship. At n=30 the interval is too wide to rank any model against another.

  4. Inference Provenance Got Harder to Log

    Ollama v0.33 adds a one-toggle gateway that surfaces local Qwen, DeepSeek, Kimi and GLM models inside Claude Desktop's model picker, running locally or through Ollama Cloud with telemetry off by default. Two engineers comparing 'the same prompt in Claude' can now be on different models, quantizations and sampler defaults with no request log. Separately, a drafted U.S. rule on China's remote access to chips would turn weight origin and host region into auditable procurement fields.

  5. A Habit Claim Its Own Statistic Contradicts

    a16z crypto argued Argentine stablecoin use crossed from crisis hedge to habit, then reported that USDC contractor-pay share and year-over-year inflation both sit at roughly one fifth of their peaks as of July 2026. A preserved ratio between two series is continued co-movement, not decoupling. The reusable asset is the April 2025 liberalization of dollar purchases: a sharp, plausibly exogenous date that belongs in your drift monitoring as a labeled regime break and works as an interrupted-time-series template.

Deep Dives

  1. The Cost Model Uber Published, and the Line It Left Out

    Unit-cost engineering at hyperscale worked exactly as designed and the invoice still multiplied, which is the arithmetic your own agent budget is walking into this quarter.

    Decompose the spend into four factors What Uber's engineering team published is an observability schema first and a finance number second. Each factor has its own owner and its own lever: session count sits with product, turns with the agent…

    3 action items

  2. Hy4's Free Weights Land on the Wrong Side of Your Self-Host Math

    Apache-2.0 at 770B reads like an exit from API pricing until the resident memory footprint, the batch collapse at long context, and an eval spread inside noise are all priced in.

    Sparse activation cuts compute and leaves memory alone The efficiency claim in Tencent's release is a FLOPs claim: roughly 49B of 770B parameters activate per token. Memory is indifferent to that. Every expert stays resident, and at FP8 that is…

    3 action items

  3. Your Factuality Number Belongs to the Endpoint You Tested

    Three independent findings put the same defect in three different layers: the surface you evaluate, the telemetry you alert on, and the kernel table quietly inflating your p99.

    The failure concentrates where the citations do The aggregate rate is not the useful number in NPR's test. On the items where Claude failed to reject a false narrative, it cited state-aligned sources more often than its peers did .…

    3 action items

The edition continues

Take the signal into the room.

Sign up or log in to read all 3 deep dives in full, plus the final take.

Read the full edition

Continue with LinkedIn