Science & Analytics

The Scientist

The Signal

RRSI treats harness search as training, rejecting wins below run-to-run variance.

The edit budget anneals the way a learning rate decays, starting with bundled edits and narrowing to single ones, which is the design choice worth borrowing if you tune agent harnesses by hand. With Claude Opus 4.8 frozen, the gain across six held-out benchmarks was +3.4 points, or +3.9 in a second account of the paper. The thing that gap doesn't tell you is which figure survives your own eval set, so plan around the smaller one.

In Play

  1. Harness search overfits its eval set

    TheSequence reports that RRSI, a harness-search method from Google Cloud AI Research, Stanford and others, gained +3.4 points on six held-out benchmarks with a frozen Claude Opus 4.8. Any loop that edits your prompts or tools against a fixed eval set ends up fitting that set, so its wins need held-out proof. Alejandro Saucedo's account puts the gain at +3.9 (43.6 vs 39.7), so check the paper before you quote either figure.

  2. Verifiers and offline indexes lift frozen models

    TheSequence reports that NVIDIA and KAIST's Mid-Harness lifted the TMAX-9B agent's Pass@1 from 50.00% to 68.03% by having GPT-5.6 Sol check candidate actions before they run. Separately, Microsoft and KAIST's CorpusMap added +6.45 RAG quality while using 0.43–0.65× the input tokens. You can try both around your current models without retraining anything. The cheaper self-verification gains, though, come to about two tasks and sit inside the noise.

  3. GPU supply now rides on credit markets

    The Information reports that Nscale is targeting a $35B IPO valuation on the strength of a $103B contracted backlog. Nearly all of that backlog depends on 12 data centers that are not yet built or fully financed. Nscale's Anthropic, Microsoft and Figure AI deals ramp in 2027–2028, so part of your future Claude and Azure capacity depends on that construction. The Financial Times reports that Nvidia is in early talks to insure lenders against losses on GPU collateral at smaller clouds.

  4. Agent tool layers become retrieval and permission problems

    Alejandro Saucedo's write-up reports that Uber's MCP platform catalogs 800+ servers and 5,000 tools, and that its scaling limit was context bloat. Uber cut the bloat by discovering tools on demand, letting agents request only the response fields they need, and writing tool outputs to files. It also registers every tool as disabled by default. In your own agents, the tokens you spend on tool schemas each turn trade directly against tool-selection accuracy. OpenAI's always-on Dots agents, with 4,000+ app connectors, widen the same permissions surface.

Deep Dives

  1. RRSI treats harness search as training, and its noise floor is the cheapest piece to copy

    Two accounts of the same paper disagree on the held-out score but agree on the mechanism your prompt optimizer is probably missing.

    Each component maps to a familiar regularizer RRSI leaves the edit space fully open. The optimizer may change any prompt, tool or control-flow step. The constraint is on how the search moves , and Alejandro Saucedo maps its three core…

    3 action items

    ●
  2. Mid-Harness's 18-point jump comes from a frontier judge, not the 9B agent checking itself

    The verifier result looks real and portable, while the cheaper self-checks sit inside the noise of a small benchmark.

    Where the 18 points come from NVIDIA and KAIST left both the TMAX-9B generator and its harness untouched on TerminalBench-Lite. They changed only what happens between proposing a terminal action and running it. TheSequence's breakdown shows how much of the…

    3 action items

    ●
  3. Your 2027 Claude and GPU capacity now rides on lenders' faith in unbuilt data centers

    Four separate financing stories point to one exposure: capacity you plan around is a forecast with an unknown chance of delivery.

    The valuation gap is a disagreement about one parameter Divide each Nscale valuation anchor by the backlog and the argument reduces to a single number. The reported $35B IPO target values the company at about 0.34x its backlog . The…

    3 action items

    ●

The edition continues

Take the signal into the room.

Sign up or log in to read all 3 deep dives in full, plus the final take.

Read the full edition

Continue with LinkedIn