Science & Analytics
The Scientist
GLM-5.3-Flash tied GPT-5.6 Terra's index despite 28% factual accuracy to Terra's 47%.
Parity holds where the work is verifiable: coding and terminal tasks, at $0.09 per task against $0.51 for the incumbent. It collapses on world knowledge. Averaging those two modes produces an aggregate that describes neither of them, so any router you key to that composite score inherits the blur instead of resolving it.
In Play
Cheap Frontier Model With a Bimodal Profile
Every efficiency figure in today's briefing was measured in a regime nobody disclosed. Z.ai confirmed that the 1M-context model topping usage charts as 'Ox Alpha' is GLM-5.3-Flash: cheap per token, with a capability split the aggregate index hides. See 'Route GLM-5.3-Flash by Task Type, Not by Index Score' below.
Ask ClaritySpeculative Decoding Decays With Concurrency
Published draft-decoding speedups collapse as concurrency rises, and no measurement reaches the widely quoted '3x'. Gate speculation on batch size before it touches your serving path, per 'The Speedup Table Everyone Quotes Was Measured at Batch Size One' below.
Ask ClarityInference Silicon Claims That Don't Reconcile
The Jalapeño ASIC claims arrive with no disclosed baseline, batch size or sequence length. The benchmark workloads are open weights, so you can reproduce the Nvidia half yourself — dissected in the speculative-decoding deep dive below. Apple separately raised Mac prices $100-200 and blamed memory costs; that pressure reaches DRAM-dependent serving next.
Ask ClarityGold Labels Expire on September 30
Amazon will shut Mechanical Turk permanently on September 30, 2026, ending a platform that once had more than 500,000 workers. A 2023 experiment found up to 46% of crowd workers used AI on tasks meant for humans, so historical MTurk gold labels may be model output now being used to score models. Requester history, HIT templates and qualification metadata become unreconstructable after the date. Export them, then measure your own contamination rate rather than arguing about theirs.
Ask ClarityGeneration Capacity Outran Validation Capacity
Uber engineers Uday Kiran Medisetty and Adam Huda disclosed that agents now author more than 70% of its pull requests, backed by 2,500 registered skills executing 20,000+ times daily and an LLM gateway carrying over 100M requests a day. Its coding agent deliberately stops at a draft PR because unvalidated agent features were consuming shared CI capacity. Reuters-seen documents show the other end of that curve: Meta's AI-heavy coding produced 405 incidents, with staff spending up to 70% of their time on remediation.
Ask Clarity
Deep Dives
- ●
Route GLM-5.3-Flash by Task Type, Not by Index Score
The 5.7x cost edge is a token price rather than a token count, and the capability profile underneath that tied aggregate splits cleanly between verifiable work and world knowledge.
The token bill says price, not efficiency GLM-5.3-Flash is MIT-licensed at $0.15/$0.50 per million input/output tokens. Artificial Analysis scores it 57 on its Intelligence Index, tied with GPT-5.6 Terra, at $0.09 per task against $0.51. The full index consumed 149M…
3 action items
- ●
The Speedup Table Everyone Quotes Was Measured at Batch Size One
Draft-model decoding, custom inference silicon and your own GPU utilization all publish their best regime, and the only figure that survives contact with production is the one you measure yourself.
Start with the arithmetic, not the tutorial Under per-token acceptance probability p and draft length K , expected accepted tokens per verification pass is (1 − p^(K+1)) / (1 − p) , and wall-clock speedup is that quantity divided by…
3 action items
- ●
Your Human Baselines Have a September 30 Expiry Date
The audit trail explaining how your oldest gold labels were produced vanishes with the platform that produced them, and the graders scoring your agents today are drifting in parallel.
What actually disappears on the date The weights are not the loss. The loss is requester history, HIT templates, qualification records and worker-quality metadata — the only archive that explains how a given label was produced, by whom, under what…
3 action items
The edition continues
Take the signal into the room.
Sign up or log in to read all 3 deep dives in full, plus the final take.
Read the full editionContinue with LinkedIn