Engineering & Technical
The Engineer
GLM-5.3 lifted Terminal-Bench from 4.6 to 28.3 without changing the base model.
DeepSWE moved the same way, 46.2 to 66.9, on GLM-5.2's base. The only lever was extended post-training and RL. That makes agent behavior a post-training artifact, free to swing that far in either direction between point releases, under a version string your deploy pipeline probably treats as noise.
In Play
Per-Agent Isolation Became Shipping Infrastructure
Containment and measurement shipped; capability merely got announced. That inverts the build order — the roadmap item you labeled blocked on a better model is almost always blocked on a gate you never built. Alibaba released OpenSandbox under Apache 2.0 and CopilotKit released OpenBot under MIT. Both give each agent its own browser, filesystem, and credentials; neither published a runtime-by-runtime cold-start number. So the one action that outranks everything below: this week, pick your highest-traffic agent path, give it its own principal and a pass/fail gate, and name an owner for both. Everything else here is backlog. The first deep dive carries the runtime, cold-start, and per-agent identity detail.
Ask ClarityDesktop Agents Are Bought Before They Are Reviewed
The Information reports Perplexity's annualized revenue moved from under $250M in January 2026 to over $750M by August, with part of the gain attributed to Perplexity Computer, an agent that drives a professional's own machine. Nothing published shows per-task success rates for it. So OS-level agents arrive expensed on a corporate card, inheriting the user's SSO session, VPN, and local cloud credentials before any security review sees them.
Ask ClarityServing Cost Became A Context-Length Problem
Netflix's GenRec scores its whole candidate set in one prefill-only pass, and compressing context from roughly 5,000 to 1,700 tokens cut serving cost to about a third. With The Information reporting Nvidia AI chip prices up about 17% on memory costs, the deep dive below has what transfers to your reranking, RAG, classification, and guardrail routes.
Ask ClarityHeadline Gains Fail Their Own Error Bars
Graft's 12-point SWE-bench Verified gap over its control carries a sampling error near 10 points (z about 1.2, p about 0.22). PlanetScale's claimed 70% drop in p95 and p99 names no workload or baseline, and Z.ai's GLM-5.3 scores are self-reported and unreplicated. The eval-harness deep dive works the arithmetic and the fix.
Ask ClarityYour Public Surface Has Unauthenticated Write Endpoints
A forged removal request got Google to de-index Muddy Waters Research's own published report on Sportradar. The firm learned of it from another researcher, Activ8Insights; both disclosures were published August 20, 2026. The deep dive enumerates the equivalent unauthenticated write paths pointed at your company.
Ask Clarity
Deep Dives
- ●
The Agent Computer Shipped Twice While Your Agents Share One Login
Two independent releases converge on the same primitives, but the value sits in per-agent identity you can build yourself — not in a v0.0.1 platform nobody has benchmarked on your image.
Why containment shipped before reliability did Z.ai moved GLM-5.3's Terminal-Bench 3.0 score from 4.6 to 28.3 , and DeepSWE from 46.2 to 66.9, on GLM-5.2's unchanged base model . Extended post-training and reinforcement learning did all of it. Two consequences…
3 action items
- ●
Netflix Deleted The Decode Loop And Context Length Became The Bill
The lift number sits inside most experiment platforms' noise band; the reusable results are single-pass scoring and speculative-decoding gains that expire on your next model upgrade.
Read the data-efficiency number, not the lift GenRec's +1.6% relative MRR is offline lift . At most companies that sits inside the noise band of the experiment platform. Online validation was four weeks on about 10% of traffic. Do not…
3 action items
- ●
A Forged Removal Request De-indexed A Published Report
Every abuse desk, registrar, and package registry pointed at your company accepts identity on assertion, and the impersonated party gets no notification — so detection has to be something you own.
Look at the call path, not the headline A third-party platform exposed a request handler that mutates your public surface area , meaning search visibility. The authentication on that handler was a plausible assertion of rights ownership. Most abuse and…
3 action items
- ●
The Long-Lead Item Is Your Eval Harness, Not The Next Model
Three of the headline gains fail their own error bars, which makes the roadmap item labeled 'blocked on a better model' almost always a missing measurement loop.
The architecture is more defensible than the number Graft's benchmark is not budgetable. The design is still worth reading. It writes a tree-sitter-derived graph of the codebase into the repository as linked markdown: deterministic operations, no API key, no network,…
3 action items
The edition continues
Take the signal into the room.
Sign up or log in to read all 4 deep dives in full, plus the final take.
Read the full editionContinue with LinkedIn