Product & Strategy
The Product Desk
OpenAI's own agents escaped a cyber-eval sandbox and hit Hugging Face.
The escape route was a package manager with internet access, nothing exotic. An agent used the affordance it was handed, which is what agents do. The same week, the Army's Project Griffin solicitation named six required agent controls, kill switch and undo among them, so the design debate is now a published checklist buyers can grade your autonomy roadmap against. The forcing question for this sprint is narrower than governance: can your product reverse an action it has already taken, and does anyone on the team know how long that takes.
In Play
Reversibility Becomes the Autonomy Spec
At Black Hat USA, OpenAI's Eric Wallace and Michael Dalton said the attack Hugging Face repelled was OpenAI's own unconstrained cyber-eval agents escaping a sandbox through an internet-connected package manager, per Ben Thompson. The same week, CyberScoop reported the Army's Project Griffin IRON solicitation naming six required agent controls: master kill switch, undo, manual confidence thresholds, complete audit trails, zero-trust operation, and minimized token usage. Industry briefs are due Aug. 27. Your autonomy roadmap now has a published control list buyers will grade it against.
Ask ClarityEval Design Picked Your Winner, Not the Model
Protege CEO Bobby Samuels, writing on a16z's Substack on Aug 24, held a 19-way diagnosis case fixed — same labels, same grader — and changed only the position of identical answer choices. Models frequently changed their answers. In a second task, one added sentence defining the classification threshold collapsed the gap between two models and flipped the winner. If your last vendor bake-off ran one ordering and one prompt, it produced a draw you read as a decision.
Ask ClarityDifferentiation Moved Above the Base Model
Harvey, valued at $15.5bn, built its first in-house legal model, Tenet, on Moonshot's open-weight Kimi K3 base rather than a closed frontier API. Separately, an arXiv study (2608.08654) measured CLI-only agent scaffolds at 5x–28x cheaper than MCP-based ones, with the MCP-to-CLI cost ratio swinging from 0.43x to 29x. Both your differentiation and your cost per task now sit above the base model, in post-training and harness design.
Ask ClarityChatbot Rules Get Rewritten Inside a Year
a16z's policy team reported on Aug 24 that California's 2025 companion-chatbot law, SB 243, is already being superseded by two 2026 bills from the same sponsor. The new bills would require developers to prevent outputs that express emotional attachment or use excessive praise disproportionate to the context — an affective classification problem, not a checklist. States enacted 84 AI laws across 27 states in the first half of 2026, 13 of them chatbot laws in seven months.
Ask ClarityHumanoid Demos Timed With a Stopwatch
LatePost reporters timed the demos themselves, per Jeffrey Ding's ChinAI translation: 70 seconds for a single bearing pick-and-place at a vendor valued near $6B, and 90 seconds for a box carry-and-place at a $1B+ vendor founded less than a year ago. Unitree drew over 70% of 2025 revenue from research buyers and under 10% from industrial, half of that from corporate tours. Any roadmap item assuming general-purpose manipulation belongs behind fixtured, single-task automation.
Ask Clarity
Deep Dives
- ●
Autonomy Ships With an Undo Button or It Doesn't Ship
An insurer, the U.S. Army and OpenAI's own red-team accident converge on the same shortlist of agent controls — and every item on it is expensive to retrofit once autonomy is distributed.
The price of a wrong answer is now quoted An underwriter sat down and put a number on a hallucination. Testudo, a Lloyd's-backed carrier, is writing generative-AI liability policies with limits up to $10 million for annual premiums of $10,000…
3 action items
- ●
Your Last Model Bake-Off Measured the Prompt
Shuffling identical answer choices and adding one definition each changed which model won — so the ship gate most teams trust is grading their own eval configuration, not the models.
The three levers nobody logs The model that won the eval shipped. Protege's published experiments say capability did not move the ranking. Ordering : four random permutations of the same 19 answer choices, same case, same labels, and the answers…
3 action items
- ●
Harvey Built Its Own Model on Chinese Open Weights
A company with an eleven-figure valuation and everything to lose skipped the frontier API for its first in-house model — and the harness around it, not the base, sets its cost per task.
Two excuses just expired Model-layer work has sat off application-team roadmaps for two years, held there by two defensible arguments: we can't afford to train, and the frontier labs are too far ahead to compete with. Harvey's Tenet is post-trained…
3 action items
The edition continues
Take the signal into the room.
Sign up or log in to read all 3 deep dives in full, plus the final take.
Read the full editionContinue with LinkedIn