Product & Strategy

The Product Desk

The Signal

Agents write 70% of Uber's pull requests while only 5.4% of models refactor cleanly.

The gap closed with 3,600 skills that each end in a verification step, not with a stronger model. Two benchmarks found the same failure shape: an agent finds a good path once and cannot repeat it. So the autonomy demo your team is scoring this quarter tells you about the first run and nothing about the hundredth. The work worth funding is the verification step.

In Play

  1. Agent Reliability Is the Binding Constraint

    TheSequence reports that only 5.4% of models tested on SWE Refactor Bench complete a whole-repository migration without breaking observable behavior. Microsoft's 507-task ThinkingBox found the same shape: agents locate a good path once and cannot repeat it. Uber says agents are attributed on more than 70% of its pull requests — the difference is 3,600 skills that each end in a verification step, not a stronger model.

  2. The AI Coworker Just Priced Itself At Zero

    Andrew Ng released OpenWorker under an MIT license with 25-plus workplace integrations and free local inference, per Simplifying AI. Ollama v0.33 registers inside Claude Desktop as a gateway provider, so one toggle swaps Qwen, DeepSeek, Kimi or GLM into Anthropic's own model picker. Tencent's Hy4 lists at $0.834 per million input tokens under Apache 2.0 weights. Capability is free; single sign-on, audit trails and per-tenant spend ceilings are not.

  3. Document Ingestion Got A Scoreboard And An 8x Cost Dial

    Datalab's Marker v2 scores 76.0% on Allen AI's olmOCR-bench across 1,403 PDFs, against MinerU's 72.7% and Docling's 50.3%. The same tool runs 2.9, 7.4 or 23.7 pages per second depending on one mode flag, and the default changes with hardware — balanced on GPU, fast on CPU. Fast mode reads equations out of the PDF text layer instead of the rendered page, so math accuracy degrades with no error in the logs. Pin the mode in config before extraction quality varies by machine.

  4. The Inference Cost Curve Is Contested

    a16z closed a $1.1B Machine Age Fund on the explicit thesis that 20-30% annual hardware supply growth cannot meet triple-digit compute demand growth, per TheSequence. The Wall Street Journal reports Nvidia has extended roughly $230B of lease and residual-value backstops, including $105B on a single OpenAI lease in Ohio. Falling list prices per token and scarcer guaranteed capacity are both real, which is why your FY27 margin model needs a flat-cost scenario next to the optimistic one.

  5. Growth That Borrowed A Tailwind

    a16z crypto measured what happens when a forcing function disappears: Argentina legalized individual dollar purchases in April 2025 and monthly inflation fell from 25.5% to 2.1%. USDC's share of contractor pay settled at roughly one fifth of its peak — not zero. The premium users pay for digital dollars compressed from over 100% to about 4%, and they still pay it. That separates willingness-to-pay for access from willingness-to-pay for convenience. Compute the same residual for your own tailwind cohorts before FY27 targets get set on a flattered base year.

Deep Dives

  1. Uber Made Agents Shippable By Instrumenting Them, Not By Trusting Them

    Two published benchmarks and one production write-up describe the same failure shape — agents find a good path once and cannot reproduce it — and only one of the three carries a fix you can copy.

    What that 5.4% is actually measuring An engineer watches an agent finish a repository migration, reruns the same task, and gets a different result. SWE Refactor Bench, from Navers Lab, Einsia.AI and Tsinghua , scores that rather than whether a…

    3 action items

  2. Andrew Ng Open-Sourced The AI Coworker; What Is Left To Sell Is Control

    A free agent, a one-toggle model swap inside Anthropic's own desktop client, and a sub-dollar frontier model landed in the same week; the governance gaps all three share are your product spec.

    The free stack runs end to end OpenWorker ships under MIT with connectors for GitHub, Slack, Jira, Notion, Linear, HubSpot, Outlook, Gmail, Google Calendar and monday.com, any MCP server, local file and terminal access, and triggers that fire on a…

    3 action items

  3. Your FY27 Margin Model Assumes Cheaper Tokens; Capital Markets Disagree

    Two streams of reporting point opposite directions on what inference will cost you, and the disagreement resolves once you separate list price from the price of guaranteed capacity.

    The contradiction, reconciled A pricing lead pulled three newsletters into one doc this week and found them arguing with each other. The Information's read is that the cost of intelligence fell again, with Tencent previewing Hy4 on August 28 and…

    3 action items

The edition continues

Take the signal into the room.

Sign up or log in to read all 3 deep dives in full, plus the final take.

Read the full edition

Continue with LinkedIn