Product & Strategy
The Product Desk
Google and Technion found models already know 95-98% of what your RAG stack retrieves.
Self-reported confidence doesn't work as a retrieval gate, because direct recall fails silently and the model sounds equally certain either way. Models still miss 26-34% of what they already hold, and extra thinking recovers 40-65% of that gap. Only 10-20% of questions genuinely need reasoning, so the routing decision you're building belongs at fact class, upstream of the model rather than inside it.
In Play
Retrieval Is Solving A Recall Problem, Not A Knowledge Problem
Google Research and Technion tested 13 models across more than 4 million responses. The models already encode 95-98% of benchmark facts, yet fail to surface 26-34% of them when asked directly. Extra inference-time thinking recovers 40-65% of those misses, and only 10-20% of facts genuinely need reasoning. If your retrieval layer mostly serves stable general facts, it is buying latency and per-call cost rather than accuracy.
Ask ClarityAnthropic's Biggest Discount Ships Switched Off
Anthropic's Fable 5.1 reads cached context 75% cheaper than Fable 5 — roughly 25% off normal use and up to 45% on agentic loops that re-read the same context every iteration. The model's five-level effort dial ships defaulted to high. Anthropic's own CursorBench numbers show 5.1 at low effort outscoring Fable 5 at high effort for a third of the cost. The saving is a config change someone on your team has to make, and the migration carries breaking parameter changes.
Ask ClarityBacklog Rubrics Are Pricing The Wrong Cost
Dharmesh Shah's September 2 essay argues AI collapsed first-order build cost to minutes while leaving maintenance and per-session user cognitive load untouched. Any RICE or value/effort score with build cost in the denominator now reads almost every feature as a quick win. In the same window, Decagon's Jesse Zhang told a16z speedrun that competitive density is a market-size signal rather than a penalty. Two of the three standard inputs in your rubric are pointing the wrong way.
Ask ClarityGoogle Won Ad Tech By Devaluing The Open Web
Judge Leonie Brinkema rejected structural remedies in the DOJ's ad tech case, per The Information's reporting, leaving Google's full stack intact. Google's winning argument was that the open web ad market is shrinking: network ads fell from 17% of its ad revenue in 2018 to 9% in Q2 2026, partly because its own AI search sends less traffic to sites. Any acquisition model compounding on Google-sourced organic traffic is compounding on an asset Google just told a federal judge is depreciating.
Ask ClarityThe Hardware Floor Under Your Margin Is Rising
ASML is raising equipment prices over the objections of TSMC, its largest customer, per The Information's chip roundup — the clearest sign that fab capex baselines are rising, not falling. Google separately contracted 400MW of Fervo geothermal power that does not deliver until 2028, with an option for 600MW more by 2030. Market-implied odds of a Fed rate hike nearly doubled to roughly 70% in a week, per Morning Brew. Model your AI feature margins on flat hardware costs for the next four quarters.
Ask Clarity
Deep Dives
- ●
The Cheapest Inference Call Is The One You Delete
Two of these cost levers sit inside decisions your team already owns, and the trap is booking either saving before you have measured it on your own traffic.
Route on fact class in code The pitched version of this optimization is elegant: a confidence score decides when to retrieve, and the bill falls out on its own. The load-bearing detail in the Google Research and Technion work kills…
3 action items
- ●
Your Rubric Is Pricing The One Cost AI Made Free
Three operators reached the same conclusion from different directions: the scarce input is no longer engineering hours, it is the discipline to say no and the nerve to ship early.
Three inputs, all pointing the wrong way A user opens the product and spends the first minute deciding which surface is the one she needs. Nothing is broken. She is paying a tax that was levied one roadmap item at…
3 action items
- ●
Google Won Its Ad Tech Case By Devaluing Your Growth Channel
One courtroom argument and one government brief moved in opposite directions for product teams, cheapening AI feature risk while putting an expiry date on search-sourced growth.
What the ruling deletes from the plan A product manager opened the divestiture line in her planning doc this week and deleted it rather than re-dating it. That was the right call. Mostly behavioral fixes were accepted and the full…
3 action items
The edition continues
Take the signal into the room.
Sign up or log in to read all 3 deep dives in full, plus the final take.
Read the full editionContinue with LinkedIn