Science & Analytics

The Scientist

The Signal

GPT-5 knows 95–98% of benchmark facts and fails to recall up to a third of them.

Google Research and Technion ran this across 13 models and 4M+ responses, so it is not one model's quirk. Extended thinking recovers 40–65% of the misses, and only 10–20% of facts genuinely need retrieval. The thing single-shot accuracy doesn't tell you is which failure you hit: knowledge the model never had, or knowledge it had and didn't surface. Those have opposite fixes, and nothing in the output signals which one in advance. Which means the retrieval budget you're sizing this quarter is probably aimed at the wrong half of the problem.

In Play

  1. Failed Elicitation, Not Missing Knowledge

    The unstaffed dispatch layer — not model choice, not index choice — is where the remaining accuracy and the remaining money now sit. First proof point: Google Research and Technion tested 13 models across more than 4 million responses. GPT-5 and Gemini-3 encode 95–98% of benchmark facts yet fail to recall 26–34% of them when asked directly, and the authors flag that models cannot signal in advance which queries will fail.

  2. Three Retrieval Failures No Embedding Swap Fixes

    ByteByteGo published a seven-mode taxonomy of retrieval failures. Three of those modes produce near-identical dense vectors by construction, which is why no embedding model upgrade touches them.

  3. A 75% Cache Cut That Inverts at 47:1

    Anthropic cut Claude Fable 5.1 cache reads 75%, from $1.00 to $0.25 per MTok. The two public measurements disagree in direction: Artificial Analysis put net cost near $3.76 per task against $3.14 for Fable 5, while Perplexity's WANDR eval found 37% lower cost at a 21% higher score.

  4. Safeguards Became an Availability Problem

    OpenAI declared Astra the first model to meet the Critical cybersecurity threshold in its Preparedness Framework. It states outright that launch safeguards will over-block legitimate work, flagging defensive security tasks and long agent runs as possible misuse.

  5. Pooled Accelerators and No Second Memory Source

    Microsoft is merging Azure and Microsoft 365 into one reporting segment because the two share AI accelerators and cloud capacity, per The Information's briefing. The previously visible 41% and 58% segment operating margins disappear into a blended number — just as Micron, Samsung and SK hynix all reach HBM4 mass production and CXMT sits at HBM3E risk production targeting 2027. Plan on flat dollars per GPU-hour.

Deep Dives

  1. The Router Nobody Built Sits Between Your Weights and Your Index

    Two independent lines of evidence point at the same missing component, and the fix starts with one new eval column plus a hard-negative set — before anyone re-embeds a corpus.

    The metric that separates "doesn't know" from "wasn't asked right" Single-shot direct-query accuracy, the number in nearly every internal accuracy report, fuses two failures with opposite fixes. Encoded-but-unrecalled means the parameters hold the fact and elicitation missed it: decoding, prompt…

    3 action items

  2. Fable 5.1's Cheapest Win Is an Effort Dial, Not a Migration

    One index point of headline capability carries a 38% price premium, and the label on the response no longer identifies which weights produced it.

    The arbitrage inside one product The Artificial Analysis Intelligence Index puts Fable 5.1 first at 66 , with Opus 5 at 63, Fable 5 at 62, GPT-5.6 Sol and Grok 4.6 at 61. The useful comparison is inside the product.…

    3 action items

  3. Safeguard Refusals Now Stop API Agent Jobs With No Human in the Loop

    Duration and persistence became safety heuristics, which makes your longest-running pipelines the likeliest to be refused — and most harnesses score a refusal as a failure.

    The enforcement asymmetry is the part to design around OpenAI's own framing is that launch safeguards for Astra will over-block legitimate work, and enforcement is not uniform across surfaces. ChatGPT and Codex users get a pause-and-review prompt. API tasks simply…

    3 action items

The edition continues

Take the signal into the room.

Sign up or log in to read all 3 deep dives in full, plus the final take.

Read the full edition

Continue with LinkedIn