Science & Analytics
The Scientist
Holding Claude Code's weekly throughput past September 13 takes roughly 20% more seats.
The temporary 50% boost gives way to a permanent 25%, and that swap is the whole arithmetic. OpenAI is separately resetting paid Codex quotas after bugs silently burned 10–50% of them, which means the capacity you think you measured last month was measured against a moving meter. Nightly eval sweeps are the workload sitting on both.
In Play
Agent Capacity Falls 17% on September 14
One thread runs through today's items: the meter moves before the price list does. Anthropic's temporary 50% boost to Claude Code weekly limits expires September 13, and a permanent 25% increase replaces it on September 14, per TLDR IT. That is an effective capacity drop of about 17% — roughly 20% more seats to hold weekly throughput. Instrument per-workflow token burn before September 13, and re-plan every sweep dated after September 14 against base × 1.25. Separately, OpenAI is resetting paid Codex and ChatGPT Work quotas after fixing token-burn bugs that silently consumed 10–50% of customer quota; nightly eval sweeps and judge-scoring runs are the exposed workloads.
Ask ClarityGrounded-QA Evals Score Answers Nobody Retrieved
Tool-call traces caught all 50 planted cheats and flagged none of 207 honest runs. So any grounded-QA or RAG score reported without those logs is an upper bound, not a measurement. The same Techpresso coverage finds frontier models surface both sides of a genuinely disputed fact only about 52% of the time.
Ask ClarityAn Adoption Stat That Counted Misconfigurations
ChinAI's Jeffrey Ding re-queried the SecurityScorecard data behind the widely repeated claim that OpenClaw use in China is "almost double" the U.S. On August 29, 2026 he found slightly more instances in the United States — from a scanner that observes only publicly reachable hosts, which is not the same thing as adoption.
Ask ClarityThroughput Is Gated by Bytes, Not FLOPs
kimi-k3-in-c ran all 2.78 trillion parameters of Kimi K3 on a single CPU, and its throughput curve is a step function. Going from 8GB to 64GB of RAM bought only 1.34x; 64GB to 128GB bought 3.54x once the working set became fully resident, per Unwind AI. Nvidia is making the same argument commercially, shipping Vera Rubin with dedicated storage and networking racks to stop flash stalling before the GPU. Profile wait-on-I/O per step before the next capacity conversation.
Ask ClarityAcquisition Method Becomes a Dataset Registry Field
Sony Music Publishing, Warner Chappell and others sued Anthropic in the Northern District of California over allegedly torrented works used to train Claude, seeking up to $150,000 per infringed work plus $25,000 per stripped copyright notice, as Pivot 5 reports. It follows the $1.5B Bartz settlement, where the court held that training was legal but piracy-based acquisition was not. Liability attaches to how a corpus was obtained — a property your ETL controls and can log.
Ask Clarity
Deep Dives
- ●
September 14 Cuts Agent Capacity While the Meter Was Miscounting
Two independent measurement failures land in the same fortnight: an entitlement change no price list records, and vendor token accounting that quietly overcharged as much as half a quota.
The entitlement was the lever Nothing in the price list moved: no dollar list price changed . Every budget model keyed to cost-per-seat now understates cost per unit of completed work , and a team without a tokens-per-completed-task denominator has…
3 action items
- ●
Your Grounded-QA Score Includes Answers the Model Never Retrieved
Trace-conditioned scoring splits one accuracy figure into three, and the drop it produces is the actual result — plus two other eval axes the research says are missing entirely.
What the trace actually buys you No new model, no new retriever. Under a function-calling or MCP-style interface the tool-call log is already being written, so the work is a join : score each answer against whether the retrieval call…
3 action items
- ●
The Adoption Stat That Inverted on a Single Re-Query
One re-pull of a public dataset flipped a headline policy claim, and the same observed-versus-claimed gap is sitting in three metrics your team reports every week.
How the number propagated The figure moved from a security vendor's scan to a news report to a think-tank analysis to downstream commentary, acquiring the phrase "China's diffusion advantage" along the way. Nobody in the chain re-queried the source. Same…
3 action items
The edition continues
Take the signal into the room.
Sign up or log in to read all 3 deep dives in full, plus the final take.
Read the full editionContinue with LinkedIn