Leadership & Executive

The Board Room

The Signal

Microsoft fixed an exploited 10.0 in Entra ID and told customers no action was required.

Remote code execution against the identity control plane is the one failure mode where tenant-side controls contribute nothing. There was nothing to patch, and no way to confirm whether a given tenant was touched in the window before disclosure. Which means the answer you hand an auditor about that period is whatever the vendor decides to publish, not anything you can produce yourself.

In Play

  1. Four Criticals, Two With No Patch You Could Apply

    The Hacker News reported that Microsoft disclosed a maximum-severity remote code execution flaw in Entra ID (CVE-2026-69836), already under exploitation, and told customers no action was required. Your identity control plane had nothing to patch and no tenant-level way to verify exposure. Three more criticals landed in the same window: five Cisco 10.0s, a GitLab flaw weaponized within days, and Rust crates with 244M downloads carrying build-time malware.

    Ask Clarity
    Try
  2. Per-Task AI Costs Are Rising While Per-Token Prices Hold

    Grok 4.6 scored 1,577 Elo against Claude Opus 5's 1,715 on Artificial Analysis' AA-Briefcase test of multi-week knowledge work — and got there in roughly half the turns and a quarter of the input tokens. For any long-horizon agent in production, the cheaper model is now the one that finishes in fewer round trips. Per-task costs are climbing even as list prices hold: Grok 4.6 gained five index points and moved from $0.36 to $0.84 per task.

    Ask Clarity
    Try
  3. Texas Halts Data Centers as Opposition Goes Cross-Partisan

    Governor Abbott says his directive halted up to 1,800 data center projects in Texas, and local opposition now runs near 75% with no meaningful partisan split. Moratorium campaigns are live in Michigan and Pennsylvania. If your 2027-2029 capacity plan still assumes 2024 permitting speed, schedule is now the binding constraint rather than chips. Bloomberg reports that Fluidstack, which builds Anthropic's data centers, now hires staff to face organized opposition before capital is committed.

    Ask Clarity
    Try
  4. The Biggest AI Buyers Are Each Other

    Meta has become one of Microsoft's largest AI customers, and Bloomberg reports AI demand remains concentrated inside the tech industry itself. Alibaba published what that dynamic costs: profit fell more than 75% while quarterly capital spending reached almost $10 billion — spending it framed as defending position rather than generating return. If your AI pipeline traces back to other tech companies' capex budgets, that is one correlated exposure wearing the costume of a diversified book.

    Ask Clarity
    Try
  5. Where Your Work Came From Becomes Checkable

    Anthropic is embedding SynthID-Text watermarks in every Claude model launched after August 2, 2026 — worldwide, not only in the EU — to satisfy Article 50 of the AI Act, with a detection API to follow. Within 12-18 months a customer, auditor, or acquirer can estimate the AI-generated share of your codebase or your consulting deliverables. Separately, Simile AI's $2B Series B puts a price on the first-party behavioral data most companies signed away in older analytics contracts.

    Ask Clarity
    Try

Deep Dives

Two of the Four Criticals Had Nothing You Could Patch

The constraint is not awareness — it is how much remediation your platform team can push through, and what your identity vendor is contractually obliged to prove.

One manifest line was the entire attack

The largest Rust crate compromise by download count needed no exploit. One dependency line went into the Cargo manifests of arrayref (244M downloads) and append-only-vec (4M), pointing at proc-macro1, a typosquat that cloned the real crate's description, author, and docs. The payload sat in a build.rs script Cargo runs automatically. The genuine upstream library code was untouched. No import required, no binary shipped: compiling was the compromise. The Hacker News reports the initial vector was a compromised maintainer account, not a code review gap, in the ecosystem most often held up as the safe one.

Two consequences get missed at leadership level. Crate deletion is not remediation, because pinned lockfiles, build caches, and vendored copies still carry the malicious versions. Any credential that lived on an affected CI runner should be treated as exposed. This is an infostealer incident at developer privilege, not a dependency bump.


Agency, not severity, should rank the next two weeks

Scanners rank by CVSS and auditors ask by CVSS, which is a reasonable way to run a queue and a poor way to run a fortnight. The dimension that gets skipped is how much power the defender actually has to fix it.

EventYour remediation agencyBlast radiusWhat actually helps
Entra ID CVE-2026-69836 (10.0, exploited)None — vendor-controlledEvery federated app and privileged workflowContract terms, independent detection, break-glass path
GitLab CVE-2026-19478 (9.4)Full, if self-managedSource code and CI/CD secretsPatch now, then decide managed vs. SLA-backed self-host
Cisco Crosswork / Secure Workload (five 10.0s)Full, but recurringSegmentation and orchestration controlsEmergency patch plus renewal-cycle leverage
Rust crates (build-time malware)Partial — deletion does not undo buildsDev laptops, CI runners, shipped artifactsLockfile sweep, runner credential rotation, egress control

On identity, the exposure is wider than one CVE. Google is tracking three Russian clusters, UNC6293 and UNC7005 as likely APT29 sub-clusters plus UNC5976, reaching mailboxes without stealing a password: app passwords, OAuth consent grants, device codes, WhatsApp device linking. MFA held every time and the mailbox was read anyway. Detection has to move from authentication to authorization, executives' personal accounts included, and those sit entirely outside company telemetry.


The bottleneck is throughput and org design

All four events were known within hours, so knowledge was never the constraint. Change-management latency and finite platform-engineering capacity are, and when four criticals contend for the same people, one quietly misses its window. GitLab was weaponized within days of disclosure. On NetScaler the sources diverge usefully: Rapid7 sees no exploitation of CVE-2026-19490, a pre-auth bypass on Gateway and AAA virtual servers, and still recommends emergency patching. A skeptic would read that as overcaution. ShadowServer counts 22,000+ ADC and 1,800+ Gateway instances internet-facing, and 22 Citrix bugs have been exploited in five years, six of them in ransomware. "Not observed exploited" is a clock, not a reprieve.

The harder problem is not security-owned, which is why it survives every tooling cycle. Build infrastructure is the highest-privilege, least-governed environment in most technology companies: owned by platform engineering, budgeted as developer productivity, monitored by nobody. A tool purchase does not close that gap. A named owner and a funded program does.

When a 10.0 in an identity provider arrives with "no customer action required," the fix that matters is not in the software. It is in the vendor terms.

One sourcing caveat: several round-ups carried truncated article bodies, so affected version ranges and prerequisites should come from primary vendor advisories before any change ticket is written.

What to do

  1. Query every CI runner, build host, and developer image for arrayref 0.3.10, append-only-vec 0.1.9, and any resolution of proc-macro1, then rotate every credential reachable from a host that compiled them.

  2. Escalate to Microsoft for written tenant-level exposure attestation on CVE-2026-69836 by month-end, and log the response — including silence — as an input to renewal and to a documented board position on single-vendor identity dependency.

  3. Fund build-time isolation and a concurrent-critical surge protocol this quarter: network-denied build scripts, pre-authorized emergency change windows, a named executive approver, and a reported percentage of criticals remediated inside the exploitation window.

Per-Token Prices Are Flat. Per-Task Costs Just Doubled.

The model that wins your workload is the one that finishes in the fewest round trips, and the vendors posting the biggest capability gains are billing you for the extra ones.

Capability is being funded by output length

The two largest capability gains this cycle were bought with tokens, not efficiency. Qwen3.8-Max added eleven points on Artificial Analysis' intelligence index and moved from $0.54 to $1.13 per task, entirely because it generates more output. Grok 4.6 added five points and went from $0.36 to $0.84. List prices per token held roughly flat through both moves. A three-year business case resting on declining inference cost of goods is therefore empirically wrong for reasoning-heavy work, and the error compounds with agent adoption rather than decaying.


The market repriced this in ten weeks

Stripe confirmed its purchase of OpenRouter at a reported $7.5 billion. The same five-day window produced Ramp buying router.com to house its own Router product, and Merge shipping Merge for Workforce with a claim of 75x fewer tokens for better results. The buyers are the tell. Neither Stripe nor Ramp is an AI lab. Both already own spend rails, and both are now treating tokens as a governed spend category sitting next to payments and corporate expense. That is a bet that the durable asset is spend attribution and distribution, not model access.

The tradeoff most org charts never named is now visible. Engineering owns latency and quality. Finance owns the invoice. Nobody owns cost per successful outcome, which is the number that decides whether gross margin survives agent volume. Vendors will monetize that gap on the buyer's behalf if the buyer declines to. The CIO Dive read of a "cautious AI era" is the same condition seen from the buyer side: deployments narrowing to measurable use cases because token consumption makes open-ended adoption hard to justify. Anthropic now publishes an enterprise guide for controlling shared usage-based spend. Category leaders do not write spend-control documentation unless buyers are being surprised by invoices.


Where the evidence is thin

A reasonable skeptic would say the magnitudes here are vendor-supplied and should be discounted accordingly. The skeptic is correct. DeepSeek priced an experimental multimodal agent model at roughly $0.87 per million words against Anthropic's ~$50, a 57x spread, on self-published figures; the model wins only 3 of 11 benchmarks by narrow margins, trails on 8, and was never run against Opus 5. Merge's 75x claim discloses no baseline and no workload. The routing coverage carries a disclosed conflict, since two of the three companies in it are the author's own portfolio. The direction is corroborated across sources. The magnitudes are not. Verification belongs before the board deck, not after it.

The winning model in 2026 is not the one that scores highest. It is the one that finishes in the fewest turns, and almost nobody is measuring that.

Read the table as a routing map, not a ranking

ModelIndex / benchmarkCost per taskWhere it belongs
Claude Opus 5Leads AA-Briefcase (1,715)$3.14 tierTop decile of task difficulty only
GPT-5.6 SolIndex 61; Terminal-Bench leader$1.23Matched on index at higher cost per task
Grok 4.6Index 61; turn efficiency$0.84Long-horizon agents; note Feb 2026 cutoff and single-vendor concentration under SpaceX
Qwen3.8-27BIndex 52; Apache 2.0Local / hosting costThe overserved bottom half of your token volume

The under-exploited line is the 27B Apache-2.0 tier, capable enough to run on a consumer machine. Counsel should clear the Max-tier license thresholds of $20M monthly revenue, 100M MAU, and $50M annual for model-as-a-service before anything ships on the licensed tiers, because that decision is cheap this quarter and expensive next. Andrew Ng's skills map, built from job postings and expert interviews, identifies disciplined evaluation and error analysis as the differentiating trait of strong AI builders. That is the same capability that makes every number above measurable in a buyer's own environment rather than in a vendor's.

What to do

  1. Commission a Return-on-Tokens baseline within 30 days: cost per successful task and turns per task across your top five AI workloads, reported to the exec team alongside cost per token.

  2. Reopen your two largest frontier-model contracts before renewal this quarter using demonstrated shiftable volume as leverage, targeting a 30% reduction in blended cost per 1,000 successful outcomes.

  3. Fund a portable evaluation harness this quarter that scores candidate models on turns-per-task against 5-8 of your own golden workflows, so a full model swap becomes a two-week decision.

Your Patent Says Five Humans. Your Press Release Said the AI.

Three moves gave outsiders a way to check where your work came from — and put a price on the one input none of them can synthesize.

The seam opposing counsel will find first

Insilico Medicine said a pulmonary fibrosis molecule was "discovered by" generative AI, then filed the patent naming five humans and no AI. Only humans can be inventors, whatever the model contributed, so the filing is correct. The cost is that both claims now sit side by side in the public record, which converts a legal necessity into a validity and credibility attack surface. Any firm that did both sits on that seam, and the count grows as generating a candidate design gets as cheap as generating text.

The remedy is weeks of unglamorous work: every AI-assisted filing from the last 24 months set beside what marketing said about the same work, and a contemporaneous human-conception log running from here forward.


Detection is arriving as an industry default, not a vendor quirk

Anthropic will embed SynthID-Text watermarks in every Claude model launched after August 2, 2026, worldwide rather than EU-only, to satisfy Article 50 of the AI Act; a detection API returning probability scores follows. Deterministic code carries fewer marks; non-deterministic code and code comments stay detectable. OpenAI, Google, Meta, and Microsoft have signed the EU Code of Practice, which makes provenance marking the default rather than one company's stance.

The board-deck version is a compliance item. The complete version is a 12-to-18-month horizon in which a customer, an auditor, or an acquirer can probabilistically estimate the AI-generated share of a codebase or a set of consulting deliverables. That reaches IP warranties, professional-services contract terms, and M&A diligence, where a probability score is enough to reprice a deal. Anthropic absorbed real backlash, including privacy objections and Steven Sinofsky invoking a "right to private thoughts free of a digital trail," which guarantees a competitor markets the no-watermark position. The answer needs to exist before a customer asks for it.


The asset side: the input nobody can synthesize

Simile AI, a roughly 60-person Stanford spinout, raised a $2B Series B from GreenOaks and Index, Fei-Fei Li and Andrej Karpathy on the cap table, for behavioral foundation models that predict what people do rather than what they say. Reported fidelity: 85% of the accuracy with which the same humans reproduce their own responses two weeks later, against 50-60% for frontier LLMs on general populations and 20-30% on niche ones. CVS runs tens of millions of simulations, Gallup is a strategic partner, Shopify built an internal equivalent called SimGym to find interventions that lift conversion.

The scarce input is not the model. It is first-party behavioral data: transaction logs, session telemetry, support transcripts. Most companies granted rights to exactly that data inside analytics and vendor contracts written before this category existed, which prices the one uncopyable input at zero.

Weak links, named: the 85% figure comes from the vendor's own research lineage, validated against a 1,000-person representative sample and not independently replicated at enterprise scale, and the post-training rests on a social-science literature with a documented replication crisis. That argues for cheap pilots with human ground truth in reserve, not a platform decision.

Provenance used to be a values debate. It is now something a buyer's counsel can measure on your codebase and your patent file.

What to do

  1. Direct counsel to reconcile every AI-assisted patent filed in the last 24 months against your public claims about AI's role, and stand up a contemporaneous human-conception log this quarter.

  2. Inventory training-rights and telemetry language across your top vendor contracts this quarter, and insert no-training clauses at each renewal.

  3. Update customer warranty language and publish an internal AI-assistance disclosure policy this quarter, before detection APIs reach buyers and auditors.

The bottom line

One pattern runs under today's items: the evidence you need to manage your own risk increasingly sits inside a counterparty's systems, and none of it arrives unless you contract for it. That breaks the assumption that choosing a reputable supplier is itself a control — reputation gives you nothing you can show an auditor, a customer, or a board. Instrumentation and attestation are now the same purchase. Name the three dependencies where you currently accept a vendor's word, and make written, tenant-specific evidence a condition of the next renewal.