Science & Analytics
The Scientist
Netflix's 0.006% GenRec lift needs about 470M users per arm under naive power math.
The test ran on about 10% of traffic for four weeks, so they are almost certainly using variance reduction they did not document. Without that description, there is nothing to calibrate your own lift estimates against. The part that transfers is cost: trimming context to about a third of its tokens cut serving cost by roughly the same share.
In Play
LLMs enter the recommendation path
Netflix's GenRec uses an LLM as the user side of a two-tower ranker. Per ByteByteGo's summary, it lifted short-term engagement 0.115% and a long-term core metric 0.006%, on about 10% of traffic over four weeks. Separately, Bloomberg reports a study in which Claude and ChatGPT steered wealthier user profiles toward pricier goods. In your recommender, whatever text you feed the LLM now acts as a set of features nobody reviewed.
Ask ClaritySelf-hosted AI servers under active attack
CyberScoop reports that PoeLLM malware has compromised more than 3,400 servers running open-source AI services since April 2026. It hides the addresses of its command servers in a poem posted on GitHub. The Information reports Anthropic's claim that Zhipu AI's open-weight GLM-5.3 matches Anthropic's original Mythos model at autonomous cyberattacks. That makes your self-hosted inference boxes both targets and possible attack tools, and network blocklists won't catch this malware's traffic.
Ask ClarityVendors reprice the context window
Simplifying AI reports that Google's Nano Banana 2.1 cuts the 1K-image price roughly in half, to $0.0336, but triples the input-token price to $1.50 per million. Nano Banana 2 is retired on October 29. Requests carrying more than about 33K input tokens now cost more than they did before. On text models, AI Breakfast's arithmetic shows Mistral Large 4's list prices ($1.36/$4.18 per million tokens) run 7–12x below GPT-6 Astra's $10/$50, not the one-third first claimed.
Ask ClarityAgent evals missing their no-agent baseline
Microsoft's SkillOpt, Renmin and Tencent's SkillAdam and OpenHands' SkillRefiner now edit agent skill files automatically. Turing Post's free excerpt gives no benchmark numbers for any of them. OpenAI and Synopsys split revenue according to how much GPT-Synopsys improves chip designs, yet TLDR Hardware found no published baseline or result. Your agent evals need a no-skill or human-only comparison arm before you let any optimizer edit your skill files.
Ask ClarityGrid access now requires curtailable compute
a16z reports that the Texas PUCT approved CyrusOne's 760 MW Freestone campus only on the condition that it can cut load or switch to backup within 30 minutes of an ERCOT call. Batch Zero's bring-your-own-power path requires cuts within one minute. ERCOT's large-load queue grew from 63 GW to 474 GW by June 2026. For your training stack, the ability to shed load on short notice is becoming a hosting term your checkpointing has to meet.
Ask Clarity
Deep Dives
- ●
LLMs now sit inside, in front of and on top of your ranker
Netflix, a wealth-steering study and agent shoppers each put an LLM at a different point in the ranking path, and each one breaks an assumption your A/B stack depends on.
The power problem behind a tiny lift ByteByteGo ran a naive power calculation on Netflix's reported numbers: two-sided α=0.05, 80% power, and an assumed coefficient of variation of about 0.33 for a binary retention-style metric. Under those assumptions, a 0.006%…
3 action items
- ●
PoeLLM and GLM-5.3 make your inference box both a target and a weapon
A botnet harvesting open-source AI servers and an open-weight model with removable cyber safeguards shift your real defenses from model refusals to host hygiene and egress rules.
Why the poem trick gets past standard controls CyberScoop describes how PoeLLM locates its command-and-control (C2) servers. It takes four words from a poem on GitHub and maps them through a hard-coded dictionary to generate server addresses at runtime. Network…
3 action items
- ●
Vendors are repricing your context window, not your outputs
Nano Banana 2.1's split price, a misreported Mistral discount and a $350B inference forecast all lead to one accounting unit that survives repricing: cost per verified task.
Where Nano Banana 2.1's discount flips Simplifying AI prices each request as the image cost plus input tokens times the input rate. Its prior prices are derived, not published: about $0.067 per 1K image and $0.50 per million input tokens.…
3 action items
The edition continues
Take the signal into the room.
Sign up or log in to read all 3 deep dives in full, plus the final take.
Read the full editionContinue with LinkedIn