Leadership & Executive

The Board Room

The Signal

One agent with the full evidence beat four-agent groups in nearly every Anthropic test.

The groups picked correctly in only 17–36% of runs. Meanwhile twelve frontier models from rival labs in the US, Europe and China returned one identical answer to the same prompt, which is roughly what overlapping training data should be expected to buy. Consensus of that kind measures correlation, not correctness, and most multi-agent roadmaps still price it as oversight, which means the review board you report to is being shown a quorum where it believes it has a check.

In Play

  1. AI Checking AI Just Failed Four Tests

    Anthropic ran the classic hidden-profile test on four-agent AI systems, and the groups picked the right answer in only 17–36% of runs, per Exponential View's analysis. One agent given the whole evidence base got it right nearly every time. Separately, twelve frontier models from rival labs in the US, Europe and China returned substantively the same answer to one identical prompt. Any control in your framework that says "ask a second model" is documenting correlated error as verification.

  2. Nvidia Reprices Two Chip Generations at Once

    Nvidia is raising prices roughly 17% on Grace Blackwell 300 and Vera Rubin 200 systems for 2027 delivery, and server OEMs have begun notifying customers, The Information reports. The increase adds at least $5 billion to one gigawatt-scale data center, implying chip systems alone already cost $29–30 billion per gigawatt. Any three-year plan built on falling cost per unit of compute is now wrong by double digits. The reporting rests on two anonymous sources and covers only some flagship systems.

  3. Engineering Leadership Is Queued to Leave

    Six in ten of about twenty CTOs and VPs of Engineering surveyed by Gergely Orosz are leaving or seriously weighing it. The stated causes are not AI itself but what AI licensed executives to demand: 20–50% cost cuts and transformation timelines nobody believes. Equity has stopped retaining them — one CTO's 2% stake pays nothing below a $210 million exit, because $110 million was raised at a 2x liquidation preference. Flatter orgs leave fewer seats to move into, so the attrition is queued, not avoided.

  4. Cheaper Tokens Are Not Buying More Demand

    Exponential View's State of AI work puts token price elasticity at 1.2–1.8, so a 10% price cut lifts usage only 12–18%. The Jevons rebound baked into most AI revenue forecasts has not arrived, which makes a price cut a transfer to the buyer rather than a volume unlock. Adoption is also narrow: since October 2023 the top 1% of firms raised AI spend per employee by $6,542 while the median firm raised it by $9.63. Software engineering leads because its output is already metered.

  5. The Labs Are Absorbing Half Your Agent Stack

    Harness-Bench ran one model across 106 identical tasks in different harnesses and scored 52.4 to 76.2 — a 23.8-point spread with no change to the weights, per Latent.Space. Roughly half of agent performance is now engineering you own rather than pretraining you rent. The catch is depreciation: trained tool use and context compaction have already migrated into model weights, and tool selection, orchestration and memory are named as next. Harness code is operating expense with a useful life, not IP.

Deep Dives

  1. Nine Agreeing Models Are Not Nine Witnesses

    Multi-agent debate, second-model checks, AI code scanning and watermark detection have all failed, and the human review layer meant to catch them is approving 93% of what it sees.

    Why agreement is not corroboration Frontier models train on heavily overlapping corpora, which means they fail in overlapping ways. Give thirty agents the same coding task and eighteen pick an identical git branch name. Blend several models' answers and roughly…

    3 action items

  2. Nvidia Deleted the Reason to Wait

    Repricing current and next generation together removes deferral as a hedge, and the unsettled question of how long accelerators last distorts unit economics by more than the price move.

    The hedge that has disappeared Nearly every three-year AI capex plan approved in the last eighteen months carries an option nobody wrote down: wait for the next generation, because cost per unit of compute falls. That option was the hedge…

    3 action items

  3. Half Your Agent Is a Harness With a Depreciation Schedule

    The same product concept produced a rounding error in 2024 and a billion-dollar business in 2025, which makes timing against the capability curve the variable your roadmap does not track.

    Same idea, four years apart, opposite outcomes Devin v1 shipped full autonomy in 2024 and returned roughly 15% task success in Answer.AI's testing. Claude Code shipped in February 2025 with a terminal, bash and file-write access and declarative permission rules,…

    3 action items

The edition continues

Take the signal into the room.

Sign up or log in to read all 3 deep dives in full, plus the final take.

Read the full edition

Continue with LinkedIn