Science & Analytics

The Scientist

The Signal

37 points of GPT-6 Astra's ARC-AGI-3 score came from the harness, not the weights.

OpenAI's 99.9% figure required an adapter preserving opaque reasoning state plus native compaction; on a standard harness, three evaluators converged at 63%, 66%, and 62.7%. The thing that gap doesn't tell you is where it transfers: the same bottlenecks yield on models already in your stack, and Trace-as-State alone moved GraphWalks from 29.2% to 81.8% on DeepSeek V4. Which makes the scaffolding line item, not the model license, the one worth arguing about this quarter.

In Play

  1. Astra's Headline Score Is a Harness Artifact

    OpenAI launched GPT-6 Astra claiming 99.9% on ARC-AGI-3. ARC Prize measured 63% on a standard harness, François Chollet 66%, and @mhmazur 62.7%. The gap came from a provider adapter that preserves opaque reasoning state, plus native compaction — so the largest lever in this release sits in infrastructure you own. Latent.Space adds the statistical reason this matters: separating a 97% from a 99% pass@1 at 80% power needs roughly 800 independently graded tasks per arm.

  2. Agent Spend Is an Observability Defect

    FrontierHarness ran coding-agent harnesses and found pass rates clustering tightly while cost varied up to 17x, with Codex highest at 66.7%. Databricks traced seven silent MCP tool bugs to an estimated $499K in wasted tokens and $1.2M in annual lost productivity, then fixed them in about an hour once tool calls were traced. Your agent bill is a tracing problem before it is a model choice. Caveat: only Codex's pass rate was disclosed, so no cost/accuracy frontier can be drawn from the published numbers.

  3. The ML Stack Enters the Actively-Exploited List

    CISA added BerriAI LiteLLM (CVE-2026-59822), JFrog Artifactory (CVE-2026-82329), Kestra OSS (CVE-2026-49869) and Starlette (CVE-2026-48710) to its actively-exploited catalog — the gateway, registry, orchestrator and serving framework of a standard ML stack. SANS separately reports litellm SSTI-to-RCE at CVSS 9.8 and three NLTK pickle and JVM-injection RCEs. Patching the gateway without rotating every provider key it holds is a half-fix, because an authentication bypass there is a credential dump.

  4. Safety Gates That Partly Measure IQ

    Ai2 ran 100 open-weight models across 16 benchmarks and found BBQ and WMDP scores track general reasoning more closely than safety. A release gate built on those numbers is partly a capability gate, so a smarter model clears the bar without being better aligned. Meta's Reels team took the opposite route: it goals its ranking models on a human 5-star/1-star panel across four declared dimensions, read over months, because roughly 1,000 concurrent experiments make per-test attribution untrustworthy.

  5. Portable Efficiency Wins on Weights You Own

    Three model-agnostic techniques shipped with measured deltas. Declarative Attention cut attended tokens during decoding 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B across 15 tasks. Trace-as-State moved GraphWalks Parents from 29.2% to 81.8% on DeepSeek V4 Pro Preview and 66.4% to 100% on GLM-5.2. SPACE action chunking cut LLM decision rounds up to 78.9%. All three attack the same bottlenecks Astra's adapter exploits, on models you already control.

Deep Dives

  1. The 37 Points Nobody Trained

    One model produced a two-thirds score and a near-ceiling score in the same week, and the regressions hiding underneath land on retrieval, scientific code, and structured workflows.

    Max effort cost less than medium effort The most useful number in the independent runs is a price. @mhmazur's ARC-AGI-3 sweeps priced the same benchmark at $26k at max effort, $38k at low, and $48k at medium . Higher effort…

    3 action items

  2. Seven Tool Bugs Cost More Than Any Model Swap Saves

    The measured variance in agent spend sits in tracing, caching, and schema strictness — and the benchmark everyone will cite for it does not publish enough to draw the frontier.

    The failure mode is a retry loop with no error to log Start with the mechanism, since it generalizes well past Databricks. A tool call fails. The failure never surfaces as an error, so the agent reads the empty or…

    3 action items

  3. Your Gateway Is Actively Exploited and Your Registry Changed Owners

    Observed exploitation, not CVSS, is the ranking that matters — and the machine holding live warehouse credentials is the same one running agents against cloned repositories.

    KEV membership is counted exploitation, not modelled severity Rank the queue by expected loss and the ordering is not ambiguous. Gitea code injection (CVE-2026-60004, KEV 2026-08-25) and JFrog Artifactory path traversal (CVE-2026-66384, KEV 2026-08-27) sit directly on the model and…

    3 action items

The edition continues

Take the signal into the room.

Sign up or log in to read all 3 deep dives in full, plus the final take.

Read the full edition

Continue with LinkedIn