Leadership & Executive

The Board Room

The Signal

Three federal agencies told AI labs to degrade output for traffic that looks like yours.

The profile the advisory flags (pooled accounts, round-the-clock traffic, chain-of-thought-heavy prompts) is also a fair description of an ordinary enterprise agent pipeline. A skeptic would say a paying customer would notice the difference and complain. Harder than that: the degradation varies across requests specifically to defeat quality evaluation, so the first sign your pipeline is being throttled will look like model drift rather than enforcement.

In Play

  1. Federal Advisory Makes Model Output Discretionary

    The NSA, FBI and CISA jointly advised US frontier labs to answer high-confidence distillation attempts with quietly degraded models, varying the degradation across requests to defeat quality evaluation, and to avoid informing affected users. Safety researchers and third-party evaluators get a carve-out. For you, paid API quality stops being a specification and becomes an unappealable classifier decision. Six Chinese firms were named, with Moonshot AI accused of distilling 17 US models.

    Ask Clarity
    Try
  2. Discovery Became a Compute Purchase

    OpenAI ran roughly 10,000 concurrent agents for about 88 hours and published a machine-checked proof that three-dimensional fluid motion can break down in finite time, spending 130 billion output tokens; MIT Technology Review notes millions of dollars of compute went into claiming a $1M prize. Its own workflow figures show inference spend rising from $14 to $600 a day to buy 3.1 agent-workdays per human shift. Any agent product priced per seat loses margin as usage grows.

    Ask Clarity
    Try
  3. Assistant Surfaces Are Revocable Leases

    Meta's Muse agent transacts against DoorDash, Etsy, Reddit, Yelp and Outlook, reaches users as a WhatsApp contact, and is free up to 100 million tokens a week with paid tiers at $20 and $100. The Information reports OpenAI simultaneously stopped accepting ChatGPT ads for image- and audio-generation products — the category its own features now serve — without updating its published policy, and Adobe learned by being cut off. Your consumer-adjacent price ceiling and your assistant distribution were both set for you.

    Ask Clarity
    Try
  4. Identity Verification Lost Its Assurance Floor

    A dark web service called Nexus was selling 153 million verified US and Canadian drivers licences with photos and home addresses — roughly 63% of all US licences. Brian Krebs verified samples against his own licence and nine others, and identity vendor IDScan has confirmed it is investigating a breach. Exfiltration ran continuously for more than a year and the data was never deleted. Every onboarding flow, account-recovery path and fraud model that treats a licence match as high-confidence signal now runs on a failed control.

    Ask Clarity
    Try
  5. Europe Now Regulates Licensing Mechanics

    The European Commission closed its SAP probe with binding commitments to abolish reinstatement fees and reduce back-maintenance fees. MLex reports regulators have begun examining Oracle's licensing practices, with third-party views being solicited while the Commission maintains no formal investigation exists. The SAP commitments now function as a published remedy template. If your EU price book carries lapse-and-return penalties or retroactive support billing, the remedy shape is knowable before any case opens.

    Ask Clarity
    Try

Deep Dives

Output Quality Is Now a Policy Decision, Not a Specification

Three federal agencies asked American labs to serve worse answers to suspected distillers without telling them, and the traffic profile they described is an ordinary enterprise agent pipeline.

The mechanic deserves a second read from procurement. The advisory tells labs to reduce reasoning depth, present correct information with different reasoning, introduce stylistic inconsistencies, and vary the degradation across requests "to complicate response quality evaluations". It also tells them not to inform the affected users. Safety researchers and third-party evaluators get an explicit carve-out and should be told.

The carve-out is the tell. Exempting evaluators concedes that degradation defeats measurement, which means a two-tier reliability regime now exists in a market where every SLA, eval score and cost-per-task model assumes a stable endpoint. Sophisticated buyers will negotiate into the top tier. Everyone else keeps paying specification prices for a discretionary output and never sees the difference, because the recommended pattern is built to survive a single-shot test.

The second front is provenance, and it arrives as a document request

The same advisory accuses Moonshot AI of distilling 17 American models, including Anthropic's Claude Fable 5, released only months earlier. Moonshot declined to comment. Its Kimi K3 is described as a genuine global hit, which is the reason the advisory exists at all. This is not litigation. It is a procurement artifact that any enterprise security reviewer, government buyer or acquirer can cite from next week, and exposure counts when it is indirect: through a multi-model router, a fine-tuned open-weight derivative, or a SaaS vendor's undisclosed inference stack.

Confidentiality moved in the same cycle. OpenAI, disputing an attribution claim, conceded it "cannot rule out that de-identified data derived from their usage of our products helped improve our models." Read that alongside model vendors pushing into finance workflows and payments, and one exposure has two faces. The supplier can vary what it delivers, and may learn from what it receives.


Where the reporting converges, and where it splits

Four independent accounts agree on the facts of the advisory and on the direction. Model supply chains are now a documented, disclosable part of enterprise risk. They split on the first move. One line of analysis says contract first: a non-degradation and disclosure clause. Another says documentation first: a model bill of materials naming every model, router and derivative in production within ten days. A third argues for a 30-day lineage sweep across production, fine-tuning and evaluation stacks, on the view that country-of-origin becomes a procurement gate before enforcement arrives.

They are sequencing the same three artifacts, and the order that holds is cheapest-first. The clause costs a letter. The bill of materials is what the first customer question demands, and producing it under deal pressure turns a memo into a discovery project. The measurement harness comes third and matters most, because longitudinal telemetry is the only evidence base that would ever support a contractual claim.

A vendor's refusal to put non-degradation in writing is itself the intelligence, and it belongs in the risk register rather than the follow-up folder.

A reasonable skeptic says no lab will ever sign this. Some won't. The negotiation is still cheapest now, before the practice normalizes and vendor legal positions harden. The predictable failure is filing all of it with the security team. This is supply-chain integrity with revenue, contract and disclosure consequences, and it needs an owner who sits at the executive table.

What to do

  1. Send every frontier model vendor a written demand this week for a non-degradation and disclosure clause: no undisclosed model substitution, a notification SLA on quality-affecting changes, and confirmation your accounts are not classified as a distillation risk.

  2. Publish a model bill of materials within ten days naming every model, router and fine-tuned derivative in production, flagging exposure to the six named firms including through vendors.

  3. Fund a continuous canary evaluation harness this quarter: a fixed golden set run hourly against every vendor endpoint, tracking correctness, reasoning depth, latency and style drift with statistical alerting.

Discovery Went on the Compute Bill; Verification Did Not Get Cheaper

The proof that made OpenAI's fluid-dynamics claim discussable was machine-checked, which turns verifiable acceptance criteria into the filter deciding which research problems agents can attack at all.

The transferable part is the architecture, not the theorem. Roughly 10,000 concurrent agents explored for about 88 hours against executable code and a cached copy of the internet, producing an analytical proof plus a Lean formalization, Lean being a checker that walks a proof line by line. A weaker model then spent 17 hours verifying it. The bill: 2.7 million messages and 130 billion output tokens on this problem, roughly 300 billion across everything attempted. Strong model explores, cheap model proves.

OpenAI then declined the $1M Millennium Prize and conceded competing priority: two mathematicians published overlapping blowup results roughly twelve hours earlier, one an Anthropic employee working independently, both using OpenAI's own coding tool. MIT Technology Review quotes NYU's Tristan Buckmaster calling it a "Deep Blue–Kasparov moment" and asking for "serious and unhurried discussion." He will not get it. The priority dispute is unresolved while the token cost of the attempt is already a line item in cloud spend.

The filter, not the milestone, changes the portfolio

The result was discussable because it was machine-checkable. A formal certificate is the one class of AI output nobody takes on faith, though a human still has to confirm the formal statement says what the field thinks it says. Innovation portfolios sort the same way. Problems with formal, simulation-based or test-based acceptance criteria are swarm-attackable at a price. Problems whose correctness is a matter of taste stay expensive, and senior human time belongs there.

Scarcity is drifting from generation to verification, and better models will not fix it. Both leading labs concede the limit: OpenAI's chief scientist Jakub Pachocki stated no lab knows how to safely reach full recursive self-improvement, and Anthropic published that autonomous successor development "has not been achieved and is not inevitable." When AI does the checking, the evaluator's weaknesses become the system's blind spots, invisible from inside the loop.


The financing side has not caught up

Cognition's most recent disclosed round cleared at roughly 53x run-rate revenue, the same multiple as its previous round after doubling both revenue and valuation, on approximately $900M ARR against up to $800M of cash burn, unaudited company figures. Nvidia, Citi, BNY and Mercedes on the roster settles the demand question and says nothing about inference margin. Capital is still paying for growth without asking, which is a reason to raise now and a warning about the repricing after.

The supply side sets the longer clock. Frontier compute carries roughly 15-month lead times and is migrating toward neoclouds: Figure's multi-billion-dollar agreement for up to 100,000 GPUs does not begin delivering until late 2027. Anything promised for 2028 rests on a commitment signed closer to now than feels comfortable, staged against demand milestones, with export-control and end-customer verification clauses so capital intensity does not outrun monetizable workload.

Compute buys discovery at a published price. Verified output is a separate purchase, and licensed data a harder one. The next dollar of differentiation goes to whichever of those a firm can actually secure.

What to do

  1. Mandate cost-per-completed-task instrumentation on every agentic workflow now, with CFO-owned kill and scale gates at named margin thresholds, before approving another agent rollout.

  2. Move agent SKUs off seat pricing to consumption or outcome pricing, piloted with two named accounts this quarter.

  3. Fund a verification layer — formal, programmatic or adversarial-eval — as a named platform capability that gates agent output before it reaches customers or production systems.

The Rule That Cut Adobe Off Was Never Published

One lab quietly narrowed who may advertise inside its assistant while another anchored agent pricing at zero, and both moves reprice every channel you do not own.

The economics are the tell. OpenAI has pitched investors aggressive advertising growth as the route to monetizing an enormous non-paying user base, then voluntarily narrowed the pool of advertisers it will take money from, excluding exactly the category its own features now serve. Companies forgo revenue to protect what they consider core. The exclusion list reads as roadmap disclosure.

The timeline deserves equal weight. Search took roughly a decade to reach credible self-preferencing complaints. App stores took about five years. This arrived inside a first monetization cycle, before the ad business is even scaled, and the restriction never appeared in the published policy. The durable damage came from the undocumented change, not the change itself. Versioned eligibility rules and a partner change-notification SLA cost about a quarter of legal and comms time.

The price of agentic work has already been set

Meta's agent holds purchase authority, interacts with email, and arrives as a contact inside an installed base measured in billions with no download step. Its free tier and its $20 and $100 paid tiers now anchor consumer-adjacent agentic pricing; Meta internally considered $200 and backed off. Acquisition models that depend on an app-open now carry a worse CAC curve than anyone distributing in-thread, and that gap widens with every tier Meta does not have to charge for.

DimensionLab-owned assistant surfaceAgent traffic on your surfaceOwned demand generation
Rule transparencyLow — change absent from published policyNone — no agent identity standard existsYours to set
Notice on revocationNo precedentNot applicableNot applicable
Leverage available to youMinimal without written warrantiesFull — you control checkoutFull
Cost profileEfficient but revocableUnmeasured todaySlower and dearer, switchable by nobody

The exposure most companies will book as fraud

Agent-originated revenue will arrive as behavior existing systems classify as abuse. Bot detection, velocity rules and CAPTCHA layers were built to stop precisely this pattern, which leaves two outcomes: blocking paying customers, or waving through automation that cannot be authenticated. Meta concedes its agent "isn't immune to attack" from prompt injection, so some of that automation will act on an attacker's instructions rather than the customer's. Human-in-the-loop confirmation on the agent's side is not a control the merchant owns; the only confirmation worth trusting is enforced server-side on irreversible actions.

Coverage diverges usefully on urgency, and the disagreement is worth naming rather than averaging. A skeptic would read the ads exclusion as administrative lag at a single lab in a single category. That reading is plausible, and it does not disturb the lesson that access is contingent. A second reading argues for capping assistant-surface dependency below 15% of pipeline outright. A third notes the agent launched US-only and 18+, a company pre-pricing impersonation and minor-protection scrutiny; if international expansion stalls on regulatory friction, the distribution threat compresses into a US-adult channel and messaging-surface urgency drops a full quarter.

Distribution you do not own is a lease, and this landlord has shown it will change the terms without telling the tenant.

The counter-position is available immediately, which makes it a this-quarter decision rather than a this-decade one. A surface that can credibly say it does not compete with its advertisers has a live wedge into creative and vertical-AI budgets that are being reassessed right now. Neutrality differentiates only while that reassessment is live.

What to do

  1. Produce the lab-surface exposure number within two weeks — the share of pipeline, signups and revenue touching a lab-owned surface — and set a board-approved ceiling against it.

  2. Decide an explicit allow, monetize or deny stance for agent traffic on every revenue-bearing surface this quarter, with server-side confirmation on irreversible actions and agent-originated conversion instrumented as its own funnel.

  3. Paper written notice periods, category-eligibility warranties and make-good remedies into every lab-surface commitment before the next spend cycle.

The bottom line

The pattern across these items is discretion without documentation: the quality you receive, the channel you rent, the assurance you inherit from a vendor's control are all governed by rules that were never published and cannot be appealed. That breaks the assumption that a paid specification is a specification, and it moves advantage to whoever can produce their own evidence instead of accepting an attestation. Fund measurement you own — endpoint telemetry, provenance documents, independent evaluation — and require every critical supplier to state in writing what it may change without telling you.