Engineering & Technical

The Engineer

The Signal

Uber's MCP gateway broke at 5,000 tools on tool schemas, not on per-call cost.

At a few hundred tokens per schema, that catalogue nears a million tokens of definitions before an agent reads one request, by an outside estimate. Each gateway call stays cheap, so request-rate dashboards never show the wall. CorpusMap's unrelated design also cut what the model sees and got better results.

In Play

  1. Uber's MCP Gateway Hit Context Limits First

    Uber's MCP gateway, fronting 800+ servers and 5,000 tools, hit context bloat before request volume. As in the compute story below, its binding limit is a resource you can't top up on demand. Instrument context tokens per turn this sprint.

  2. A Stronger Verifier Lifted a 9B Agent 18 Points

    NVIDIA and KAIST's Mid-Harness lifted a 9B agent from 50.00% to 68.03% Pass@1 on TerminalBench-Lite by having a stronger model vet candidate actions. Self-verification reached only 54.76%, per TheSequence.

  3. Compute You Can't Buy on Demand

    Gergely Orosz reports CPU shortages. The Information reports Nscale's $103B backlog rests on 12 data centers not yet built or fully financed. The FT reports Nvidia is weighing insuring GPU-cloud lenders. Together they undercut the assumption behind both your autoscaler and your future-dated GPU reservations: that capacity arrives on request.

  4. An Anthropic Engineer's Report Moves AI Incident Response to 'Maybe'

    Anthropic engineer Alex Palcuie reported real problems using Claude for incident response, a trajectory Sylvain Kalache summarized as 'from no to maybe.' Separately, Lorin Hochstein analyzed an incident report that blamed both correctly following an incorrect procedure and incorrectly following a correct one. For your on-call tooling, automation removes execution slips but runs a wrong runbook faithfully and fast. A 'maybe' still leaves autonomous remediation unproven.

Deep Dives

  1. Uber's Gateway and CorpusMap Both Shrink What the Model Sees

    Two unrelated designs cut agent input and got better results, which makes your context budget, not your request rate, the first number to instrument.

    Why the catalogue broke before the traffic The ML Engineer's back-of-envelope math explains where the wall landed. A tool schema is a few hundred tokens, so Uber's catalogue costs an agent roughly a million tokens of definitions before any request.…

    3 action items

    ●
  2. A Frontier Judge Carried the 9B Agent, and Self-Checking Didn't

    The ablation locates reliability in an independent, stronger model reviewing short candidates. That is a cost you pay per step, so spend it only where actions can't be undone.

    The ablation is the result Mid-Harness leaves the generator and harness alone. Each step samples candidate terminal actions, N=8 in the headline run, and a verifier vets them before anything executes. TheSequence's ablation is the part worth keeping. Distilling the…

    2 action items

    ●
  3. Your Autoscaler and Your 2027 GPU Reservations Share One Bad Assumption

    CPU shortages and unfinanced data center sites break the premise that capacity shows up when requested. That turns fleet efficiency and delivery evidence into the real supply plan.

    A backlog is a dependency chain Remove the valuation and Nscale's contracted capacity is a sequence of steps, each of which can fail on its own. Financing closes, the site gets built, power is energized, chips arrive, racks burn in,…

    3 action items

    ●

The edition continues

Take the signal into the room.

Sign up or log in to read all 3 deep dives in full, plus the final take.

Read the full edition

Continue with LinkedIn