Leadership & Executive

The Board Room

The Signal

Instinct's ask jumped from $2.5B to $10B in a month, with investors now paying in GPUs.

Instinct went from $2.5B in August to a roughly $10B ask in September, with a product that still isn't widely available. Capital is commoditized and capacity is not, so a rival who can allocate chips outbids your cash on the same deal, and SpaceX re-scoping its build leaves 2027 supply forecasts unstable.

In Play

  1. Passkey Enrollment Is Now the Phishing Target

    One pattern runs through today's items: the engine worked and the wrapper around it failed. Start with the live one — attackers are phishing passkey enrollment rather than passkeys, using fake IT helpdesk personas to walk employees through OAuth device-code flows that issue Microsoft 365 refresh tokens, per CSO's reporting. Deep dive below on why the fix is policy, not procurement.

    Ask Clarity
    Try
  2. Agent Write-Actions Failed at the Payment Layer

    Meta's Muse booked two hotel rooms when asked for one, then reported the card uncharged when it had been. The $408 came back on a Marriott clerk's discretion, not a product safeguard. Deep dive below on the transaction layer no vendor owns.

    Ask Clarity
    Try
  3. ServiceNow Concedes Seat Pricing, Annexes Security

    ServiceNow is shifting to consumption-based monetization while mounting a cybersecurity expansion play, per CSO's reporting. Deep dive below on what that does to per-seat benchmarks and to your security-operations budget line.

    Ask Clarity
    Try
  4. Capability Claims Arrive as Valuation Instruments

    Three frontier claims in 48 hours: an unnamed OpenAI model on forced Navier–Stokes, then "substantial progress" on a second problem told to the New York Times, then a rumor placing Anthropic on a third. Deep dive below on the disclosure pattern and what it means for roadmap evidence.

    Ask Clarity
    Try
  5. Investors Have Started Paying in GPUs

    Investors are now bundling GPU allocation with cash, per The Information. Anjney Midha pairs chips with checks. Instinct — founded in 2025 and valued at $2.5B in August — is seeking capital at roughly $10B in September while describing a profound compute shortage and remaining not widely available. Capital has been commoditized; capacity has not. Your corp-dev currency is weaker than a compute-backed rival's, and SpaceX re-scoping its own data center build means the supply forecasts under your 2027 cost model are unstable.

    Ask Clarity
    Try

Deep Dives

Your Passwordless Program Bought a Control, Not a Process

The costly assumption is remediation: issued tokens outlive the credential reset your incident playbook is built on, and two of the four exposure windows sit on a vendor's clock.

One pattern, four windows

The exposure window is the unit of attacker economics now, and today's reporting prices it four ways. A Chromium fix landing upstream before Chrome ships it downstream is a scheduled, reusable window. ConnectWise's ScreenConnect authentication bypass left privileged remote-access tooling exposed across downstream customers for five days, entirely on the vendor's clock. Google's Early Access channel puts over-permissioned utilities onto employee Android devices with reduced vetting. A passwordless rollout creates months in which users have no mental model of what a legitimate prompt looks like, and attackers are pricing that confusion accurately.

VectorWhat actually changedWindow lengthWho controls the clock
Enrollment abuse via helpdesk pretextAttack moved from credential theft to token issuanceDuration of your entire rolloutYou
Upstream-to-downstream browser patch gapPublic fix precedes shipped fix, predictablyEvery Chromium release cycleVendor
Remote-access auth bypassPrivileged tooling exposed pre-patchFive daysVendor
Pre-release mobile app channelReduced vetting reaches managed and BYOD devicesPersistent until policy closes itYou

The remediation assumption that fails

Device-code compromise issues refresh tokens that survive password resets. Rotation invalidates the password and leaves the refresh token live, so a playbook centered on credential rotation closes the incident on paper only. A tabletop exercise surfaces this in an afternoon. Eviction time is the metric to instrument: detection to dead token, target under an hour.

Where this stops being a security problem and becomes a procurement one

Two of the four windows close with policy. Two are hostage to a supplier's release cadence, and patch latency belongs in contracts for that reason, not only in engineering retrospectives. Few suppliers will sign a measured disclosure-to-patch number, and the ones that will are the ones worth the renewal. Written into top vendor agreements and into third-party risk scoring at renewal, that number converts an engineering complaint into an enforceable term. Suppliers will begin quoting their own patch latency in sales material, and the buying side should have its own figure ready when they do.

What the board pack should stop reporting

Both security desks reached the same structural conclusion, and one states it plainly: confidence and preparedness are not necessarily true measures of security. That is a candid concession from a genre that sells readiness, and it marks the vendor-sponsored readiness report as a weak buying input. The same skepticism applies to internal reporting. Control coverage, attack-path reduction, and detection and containment latency are measurable. Self-reported readiness scores are what most board packs carry instead.

Sourcing caveat: the reporting is headline-level, with no CVE identifiers, affected versions or indicators of compromise. The strategic pattern holds across both sources; the technical particulars need validation against vendor advisories before anyone briefs detail.

The attacker enrolls a passkey during the rollout window. Resetting the password leaves the enrollment intact.

What to do

  1. Gate or disable OAuth device-code flow tenant-wide, then run a token-revocation tabletop with a one-hour eviction target.

  2. Add measured disclosure-to-patch commitments to your top vendor contracts and third-party risk scoring at the next renewal cycle.

  3. Replace self-reported readiness metrics in board security reporting with control coverage, attack-path reduction and containment latency before the next board cycle.

The Agent Layer Nobody Owns Is the One That Moves Money

Meta brought the largest consumer distribution asset on earth to this fight and still lost to its login screen, which tells you exactly where the defensible position sits.

The onboarding stack is the damning part

The booking error is the quotable failure. The onboarding stack is the strategically important one. The Information's reviewer could not complete cross-device SMS verification — the codes errored out, forcing him to create a second account — and payment setup routed him through both Chase and Stripe before the agent could transact at all. At launch, the app sat buried in the App Store beneath unrelated apps sharing its name. Meta owns the largest consumer distribution asset on earth and still could not deliver a clean first session.

The tasks that worked are worse news than the one that failed. Muse did complete a Resy reservation and schedule an Uber, and the reviewer could not honestly say either was faster or easier than opening those apps directly. The benchmark for an agent is not another agent. It is the incumbent vertical app — which starts with the user's identity, payment method and live session already in hand.


Where value migrates when trust is the constraint

SurfaceOnboarding frictionIdentity and payment contextDefensibility
OS-level agent (Apple, Google)Near zero — device already knows the userNativeVery high; cannot be displaced head-on
Assistant-embedded (ChatGPT, Claude, Gemini)Low — existing account and habitPartialHigh on distribution, unproven on transactions
Vertical app-embedded (Uber, Resy)Low — user is already in-sessionNative to the verticalModerate; strongest as an action provider
Standalone agent appHigh — auth, payment and context built from scratchNone at startLow; pays for what incumbents get free

The position nobody has taken

No one owns agent transaction integrity: deduplication, charge auditing, receipt reconciliation, dispute automation, and a liability model for who eats a duplicate booking. Every OS-level and assistant-level agent will need it. None of them will want to hold the liability. That is an infrastructure position with a defined buyer set, and it was surfaced by a failure rather than a pitch deck — the cheapest way to find one.

The quiet risk your dashboard will show late

Demand depth. The reviewer's list of tasks that naturally occurred to him numbered a handful, and then he ran out. If personal agents have a discovery problem rather than a capability problem, it surfaces as a retention cliff roughly two quarters after launch — long after the roadmap is committed and the team is hired.


Where the two reads converge

Set this against the frontier research picture from the same week and the diagnosis agrees: the model was never the limiting factor. OpenAI's Navier–Stokes run was process innovation wrapped around a frozen model. Muse's failure was process absence wrapped around a competent one. Capability is not what separates the winners in this category.

Caveat worth holding: this is one reporter's few days with a launch-week product, not a cohort study, and the specific bugs will be patched. The structural argument — that agents get absorbed into operating systems and existing apps rather than winning as destinations — does not depend on the bug count.

The winning agent surface is embedded, and the winning capability is a purchase that never happens twice.

What to do

  1. Freeze and red-team every agentic write-action in production and staging within two weeks: idempotency keys, post-action verification against the payment system, and a human confirmation gate above a named dollar threshold.

  2. Make the build-versus-embed call on any standalone agent app this quarter, and redirect the budget into an agent-callable API on surfaces you already own.

  3. Commission a build-versus-buy memo on the agent transaction-trust layer this quarter, with two or three named targets and TAM sized against agent-initiated transaction volume.

An Incumbent Just Repriced Itself Away From Seats and Into Your Category

The trade is pricing power for forecast volatility, and the second move puts a workflow platform with deep CIO relationships inside the security-operations budget line.

The trade hiding inside the pricing change

Consumption pricing looks like alignment — you charge for value delivered. What it actually does is swap pricing power for forecast volatility, and revenue your CFO cannot predict is revenue your board discounts. That is a revenue-quality decision, not a packaging decision, and it is why most teams defer it until a large customer forces the conversation mid-renewal. The option value of having modeled it first is the entire play.

Why AI forces the question at all: a seat prices access to work. Agents do the work without occupying a seat. Once agent-mediated usage substitutes for logins, the metering unit stops growing when the value grows — and the gap widens quietly, inside accounts that look healthy on a seat dashboard.

DimensionSeat-basedConsumption-based
Revenue predictabilityHigh — the CFO's preferenceLower; variance by segment
Expansion mechanicHeadcount growthUsage growth, including agent traffic
AI exposureDirect — automation compresses seatsHedged; automation can raise usage
Buyer objection"We're paying for logins nobody uses""We can't budget an unpredictable bill"

The second move is the one aimed at you

Security is now the default adjacency for any software incumbent hunting AI-era growth. Practically, that means budget consolidation pressure on the security-operations line, delivered through a CIO relationship you probably do not own. Workflow adjacency is a genuine advantage in procurement and a genuine weakness in detection depth.

Your counter-position is narrow and real: security-native depth, detection-engineering credibility, and pricing predictability sold against consumption-bill anxiety — the same anxiety the pricing pivot creates. That last one is the underrated argument, because it turns the incumbent's two moves against each other. The window has a shelf life measured in quarters, and it only works if you have articulated the position before a deal is contested.

Where the reads converge, and where they part

Two independent security desks flagged the same two moves in the same week, and both treated them as market structure rather than product news. That convergence is the signal worth acting on. They diverge on emphasis: one frames security as the growth hedge of an incumbent whose core metering unit is impaired, the other as consolidation pressure landing on your buyer's budget. Both readings put the identical decision on your desk — partner or compete — and both agree that reacting after the first lost account is the expensive path.

Neither source published contract terms, pricing tiers, or a timetable. Treat the direction as solid and the mechanics as unconfirmed until the company details them.

When an incumbent changes how it charges, it is telling you the unit you both sell by is broken.

What to do

  1. Put a consumption-pricing scenario in front of the board this quarter — a model, not a decision — with NRR sensitivity by segment and a three-year revenue band under seat compression.

  2. Take a documented partner-or-compete position on the security expansion within 30 days, and instrument win/loss against them starting with the next contested deal.

Three Frontier Claims in 48 Hours, and No Named Model Behind Any of Them

Mathematics fell first because proofs are machine-checkable, which makes the automation queue in your own business auditable this quarter and an organized professional counterparty inevitable.

Why mathematics went first

Not because it is hard. Because it is legible to a reward function — proofs are machine-checkable, so reinforcement learning can be aimed at them cleanly, and a single problem can sustain a forty-year career. Verifiable output plus deep human expertise is the selection rule for what automates next. That rule is auditable, so you can run it against your own P&L this quarter rather than waiting to be surprised: score your top workflows on oracle clarity (is the output machine-checkable?), data availability, and time-to-parity, then rank by time-to-commoditization.

What the headline result actually cost to produce

Roughly 10,000 agents coordinating for 88 hours — sharing findings, using tools, working one problem. That is not a better model. It is an industrial process wrapped around a frozen one, which carries two consequences for you. Competitive advantage in reasoning has migrated from model weights to inference orchestration economics. And your cost model needs to move from cost-per-token to cost-per-resolved-task, with a scenario band for 10x–100x variable inference on hard problems. Margin structure, not model quality, is what a 10,000-agent norm threatens.


The pipeline with no verification step in it

Within 48 hours, the swarm result had been repackaged into a $1,495 conference pitch, carrying a second-hand Sam Altman quote that he "did not expect a result of this magnitude to happen so soon." Frontier claim to commercial narrative to somebody's strategy deck, in days, with no verification anywhere in the chain. And the Navier–Stokes result came from a model nobody can name, disclosed through a newspaper.

The cost of being 72 hours late to a real signal is trivial. The cost of committing capital against a retracted one is not.

The counterparty forming on the other side

Twenty-five Fields medallists published "A Severe Misalignment of AI in Mathematics," warning that mass production of true/false statements at accelerating pace "could destroy fertile ground instead of breathing life into new ideas." Their ask was procedural — paced collaboration. The industry's answer was to diagnose a wounded ego. That is how a technical debate hardens into a political coalition, and medicine, law and engineering have the same institutional machinery available. Opening pacing and attribution conversations with the professional bodies in your core verticals is cheap insurance now and expensive lawfare later.

The org-design warning underneath it

William Thurston became so dominant in foliations that advisors steered graduate students away; the field died of his success and recovered only after he left. AI does not leave, tire, or retire. Applied inside your company: automate the junior tier of an expertise function and you dissolve the path that manufactures your senior tier — a seven-to-ten-year lag with no fast remedy. The leading indicator is already visible in schools, where students post better measured performance on weaker underlying competence. Measure competence separately from output, and put senior-bench depth on the board dashboard.

Where the reads diverge: one treats this run as proof that inference orchestration is the new moat; another treats it as a valuation instrument timed to an IPO window. Both agree the weights are not the differentiator. The Anthropic third-problem claim remains a rumor.

What to do

  1. Score your top 20 revenue-generating workflows on oracle clarity, data availability and time-to-parity this quarter, and rank them by time-to-commoditization before next year's plan locks.

  2. Institute an evidence gate now: no vendor capability claim enters roadmap or procurement until it clears your own eval harness on your own data.

  3. Re-baseline the inference cost model from cost-per-token to cost-per-resolved-task this quarter, with a 10x–100x elasticity band for hard tasks.

The bottom line

In every failure here, the engine worked and the wrapper around it did not. The reasoning was sound, the cryptography held, the proofs checked out — and the layer that verifies, reconciles and stands behind the result is the layer nobody has built, priced, or assigned an owner. That breaks the assumption most plans still rest on, that buying capability buys accountability. From here the margin belongs to whoever can stand behind an outcome rather than produce one. Name a single executive accountable for verification across agent actions, credential enrollment and vendor claims this quarter, and fund it as product infrastructure, not quality assurance.