Investment & Market Intelligence

The Investor

The Signal

The Big 3 AI labs earn 8x per token on 52% of volume.

The premium everyone has been treating as a moat is a price level, which is a less comfortable thing to own. Cursor's router matched frontier output quality at roughly 60% lower cost, and OpenRouter is now fielding multibillion-dollar takeover interest, which suggests someone with a balance sheet agrees that routing is where the margin went. This is probably too early to call, but anything in the book whose margin or revenue passes through frontier per-token pricing deserves re-underwriting now rather than after the next price cut.

In Play

  1. The Model-Layer Rent Is Being Routed Away

    Vercel's AI Gateway data, surfaced by Exponential View, shows OpenAI, Anthropic and Google capturing 90% of spend while serving 52% of tokens — roughly 8x the revenue per token everyone else earns. Cursor's router reproduces frontier-perceived quality at about 60% lower cost, and The Information reports OpenRouter fielding multibillion-dollar takeover interest. Every position whose gross margin runs through frontier API pricing needs a sensitivity table before its next mark.

    Ask Clarity
    Try
  2. AI App-Layer Multiples Inverted

    a16z disclosed marks that invert the usual pattern: Harvey sits at $11B on hundreds of millions of ARR (~25–45x), while Decagon is $4.5B on an eight-figure base after 18 months (~50–150x), having tripled in under six months. The same piece describes founders discounting steeply — and sometimes paying customers to adopt — for logo revenue that 'barely registers.' At 50–150x, a 30% error in revenue quality is a catastrophic pricing error in your comp tables.

    Ask Clarity
    Try
  3. Nvidia Moves From Supplier to Royalty Holder

    The Information reports Nvidia will take a cut of some customers' cloud revenues — converting a hardware sale into a perpetual claim on downstream revenue. In the same cycle, The Information Briefing notes Nvidia fell 5% on a $500B SK partnership that was substantively a re-run of early-June announcements and rests on letters of intent only. The reaction function flipped: the market now charges for AI capex promises instead of paying for them.

    Ask Clarity
    Try
  4. General Scaling Erases Two Claimed Moats

    Import AI reports Claude Opus 4.7 completing a MirrorCode task in 14 hours for $251 of inference — work Epoch AI and METR estimate at 2–17 human weeks. Separately, Anthropic's Project Fetch had Opus 4.7 finish quadruped robot tasks autonomously in 9 minutes 35 seconds against a 181-minute human-plus-AI record, with Anthropic stating this was not the result of any robotics-specific effort. Two diligence templates break: code as a moat, and proprietary teleoperation hours as a moat.

    Ask Clarity
  5. Security Capital Rotates From Prevention to Access Architecture

    Verizon's 2026 DBIR ties ransomware to 48% of breaches, and CSO reporting has all four leading perimeter vendors — Palo Alto, Fortinet, Citrix and Check Point — under active campaigns, with a CVSS 9.3 Check Point SmartConsole flaw granting unauthenticated admin. CyberScoop reports Senator Wyden asking CISA for a binding two-year deadline to eliminate internet-facing legacy federal VPNs. Appliance-renewal durability and mandate-dependent compliance ARR both deserve a haircut this quarter.

    Ask Clarity

Deep Dives

The Routing Layer Just Got Its Price Discovery Event

Consensus called model routers thin wrappers with no defensibility; two independent cost proofs and one multibillion-dollar bid say the decision right is the asset.

What a router actually owns

A router owns no model. It owns the model-selection decision, which is where the spread between token volume and token spend gets captured, and along the way it accumulates something no single lab can replicate: a per-model, per-task performance table. Devshot's reporting puts numbers on how valuable that table is. Matching edit format to model swings agent success from 66% (DeepSeek, unified diff) to 94% (Doubao, JSON Patch). Reviewer pairing is asymmetric in the same direction. Claude checking Codex lifts pass rates from 71.6% to 89.7%, while Codex checking Claude drops accuracy from 91.4% to 82.8%. Those are not tuning footnotes. That is the routing table, and routing tables compound with usage.

Two independent cost proofs, one bid

The substitution here is measured rather than theoretical, which in this market is rarer than it sounds. Cursor's Auto Intelligence mode produces output users judge as good as the premium model at roughly 60% lower cost, per Exponential View's read of Vercel gateway data. Cursor separately cut a browser-building experiment from $10,000 to $1,300, about 87%, by routing routine work to cheap execution models behind a frontier planner. And Lenny's Newsletter reports a blind seven-model practitioner benchmark in which Claude Sonnet 5 scored 77 against Opus 5's 78, which is a vendor's own mid-tier model within one point of its flagship.

Then the price discovery. The Information reports OpenRouter fielding multibillion-dollar takeover interest. Strategic acquirers are paying for demand aggregation across models, cross-model performance data, and the switching friction that appears once a router sits in the critical path of production inference. The market spent two years dismissing that as a moat argument.


Where the sources genuinely disagree

The deflation story is not clean, and the disagreement is the useful part. Anthropic shipped Opus 5 at unchanged pricing of $5 per million input and $25 per million output tokens while claiming efficiency gains, passed none of them through as a price cut, and opened a paid latency tier at roughly 2.5x speed for 2x the rate. That is price discrimination, not a price war. ChinAI reports Moonshot going the other way entirely: Kimi K3 launched at $2.30 per million blended tokens, a 3.5x output-price increase over its own predecessor, thirteen times DeepSeek V4 Pro's $0.18.

If frontier prices stop falling, the automatic COGS tailwind penciled into your AI application models disappears — and the only remaining path to software-grade gross margin is architectural.

Two readings survive, and this is probably wrong, but the first is the one to underwrite. Either capability parity is converting into pricing power at the model layer, in which case app-layer margin expansion has to be engineered rather than waited for. Or Moonshot's increase is a capacity constraint wearing a strategy costume, since it suspended subscriptions within 48 hours of launch because demand exceeded compute. Allocation looks identical under both. Durable margin sits with whoever decides which model runs, not whoever trained it.

The reflex to avoid

The cheap version of this trade funds another gateway. The expensive version keeps crediting 'we use the best model' as defensibility while the routing question goes unasked in diligence. Founders who have already tested output parity at 60% traffic diversion are structurally cheaper to scale, and the re-architecture money not spent on them later is money available for the next position. Founders who have not have just told you something about their rigor.

What to do

  1. Commission a routing-adjusted gross-margin case on every inference-exposed position this month, modelling a premium collapse from 8x to 3x revenue per token and 50%+ of volume diverted to cheap execution models.

  2. Add one gating diligence question to every AI application memo: what is gross margin if 60% of traffic routes to non-frontier models, and has output parity been tested with human scoring?

  3. Map five routing, model-selection and eval-observability teams for first meetings this quarter, prioritising those holding proprietary per-model performance data over integration surface.

Harvey at $11B, Decagon at $4.5B: The Multiple Went Backwards

The AI app layer is paying a fatter multiple for logo velocity than for revenue base — which makes revenue quality, not TAM, the variable most likely to break a mark.

Read the marks as a statement of what is being bought

a16z published a go-to-market framework, which is fine, but the investable content is the valuation data stapled to it. Harvey: $11B on hundreds of millions of ARR, call it 25–45x. Decagon: $4.5B on an eight-figure base built in 18 months, call it 50–150x, after tripling in under six months on the back of 100+ new enterprise customers in 2025. The company with less revenue carries the richer multiple. What is being priced is not the revenue base. It is logo velocity as a proxy for future net revenue retention.

That is precisely the bet a16z's own failure list warns about, which is a pleasant sort of honesty. 'Dying of indigestion', waking up with 200 customers and 50 underwater, is the specific failure mode of a velocity multiple. The non-linearity is what matters for recovery scenarios: 50 unhappy customers is churn, 500 is a reputation problem in a market where references travel faster than renewals.


The revenue-quality problem is described from inside the category

The uncomfortable disclosure sits mid-post, where such things usually sit. Founders are burning through rounds courting first Fortune 100 customers, offering steep discounts and in some cases paying customers to adopt, for revenue that 'barely registers.' Roughly 500 marquee accounts are being pitched by effectively every AI startup, and those accounts extract concessions because they can see the seller is desperate. If that is happening at scale, a meaningful slice of the ARR in the pipeline is discounted, pilot-stage or non-recurring revenue wearing an ARR label.

At 50–150x revenue, a 30% error in revenue quality eats the entire return.

The saturation risk on the other playbook

The lighthouse cohort has the mirror problem. Hebbia is already past 40% of the largest asset managers by AUM, including KKR and BlackRock. Deep penetration of a concentrated, status-legible market is exactly what makes those logos travel, and it also means over half the addressable cohort is already consumed. The counter-thesis is that those references open the mid-market at near-zero cost, and that is genuinely possible. A terminal-growth assumption on a lighthouse asset still requires an explicit down-market or international plan with unit economics attached, not a TAM slide.

The same tell shows up one layer over. The Information reports Mercor's growth concentrated in the largest AI labs, growth funded by a handful of counterparties who could each build the capability in-house. That is a services business wearing SaaS clothing. It deserves a customer-concentration haircut and an insourcing scenario, not a software multiple.

Where the competition is thinnest

The framework's most useful output for sourcing is geographic rather than strategic. Against those 500 contested accounts sit roughly 50,000 companies outside normal venture networks, or rather the more interesting version of that number: the manufacturers and distributors in Ohio and Texas nobody bothers to call. Michigan is on the list as well, and no associate is flying there this quarter. That opacity is the mispricing, and most sourcing funnels structurally cannot see it. Caveat worth holding: every positive case study in the piece is a16z-affiliated, and the operating metrics cited are vendor-supplied and unaudited.

What to do

  1. Require ACV distribution, discount-to-list, services-versus-recurring split, top-five customer concentration and pilot-to-paid conversion in every AI app-layer data room going forward.

  2. Re-underwrite every late-stage lighthouse asset in the book this quarter against an explicit cohort-saturation case, requiring a costed down-market or international expansion plan before crediting terminal growth.

$251 of Inference and a 9-Minute Robot: Two Diligence Templates Just Broke

The failures in the benchmark data, not the successes, are the usable screen — and they say defensibility now lives in dense implicit specification.

Start with the eight targets nobody finished

Import AI reports Claude Opus 4.7 completing a MirrorCode target in 14 hours for $251 of inference, against Epoch AI and METR estimates of 2–17 human weeks, or roughly $11.5k–$98k of fully-loaded senior engineering labour. The benchmark asks for reimplementation of real software — Apple's 61,000-line pkl configuration language among them — from command-line access alone, no source code, no web.

Everyone will quote the arbitrage. The screen, or rather the more interesting version of the screen, sits in the residual. Eight of 25 targets were never solved to 100%, and the failures cluster diagnostically: ruff (Python linter), giac_subset (mathematics), mailauth (email authentication). What they share is dense implicit specification — correctness defined by thousands of undocumented conventions and standards interactions that black-box observation will not surrender.

Which converts neatly into a grading rule for app-layer software. High specification density (tax and regulatory compliance, clinical protocol, payments and settlement rails, security primitives, standards-heavy interoperability) keeps its moat, because correctness is unobservable from outside. Low density does not: CRUD workflow tools, thin orchestration, single-function utilities. If a model can probe the API, it can rebuild the product for the price of a lunch.


The robotics moat moved, and two independent sources say so

Anthropic's Project Fetch Phase Two is the second break. Opus 4.1 could not perform the quadruped tasks at all in August 2025; by May 2026 Opus 4.7 completed all but one autonomously in 9 minutes 35 seconds against a 181-minute human-plus-AI record. Anthropic's own framing is the load-bearing part: 'this progress is not the result of a concerted effort to improve the robotics capabilities of our models.' Sunday Robotics reports the same mechanism from the startup side — scale pretraining, then hill-climb with minimal in-house data — at 99.1% success across 778 folds and nine garment types, with a family beta committed for Fall 2026.

Two independent confirmations of one causal claim is a thesis, and the thesis contradicts the most common robotics moat slide in venture: proprietary teleoperation hours at volume.

What this changes in the diligence template

Retire 'teleoperation data hours owned' in favour of base-model access, post-training loop velocity, and a real-world recovery-data flywheel, then re-score the robotics deals already passed on the old criterion. The opportunity cost is the point: capital committed to buying teleoperation volume is capital not committed to the loop. And note the overhang, which sits on every robotics cap table whether or not anyone discloses it — frontier labs are now latent robotics platform players with zero robotics capex.

Two caveats belong in the memo, and this thesis is probably wrong in at least one of them. Demo tasks are not commercial reliability, and MirrorCode is flagged by its own authors as 'perhaps a little too easy' with 17 of 25 targets achieving perfect runs at release. Sunday concedes the residual gap is long-tail failures only visible after repeated real-world runs. The Fall beta will show them or it will not.

What to do

  1. Grade every app-layer deal in the active pipeline high, medium or low on specification density before the next investment committee, using the unsolved benchmark cluster as the reference for what remains hard.

  2. Rewrite the robotics diligence template this quarter to score base-model access and post-training loop velocity, and re-score deals previously passed on teleoperation-data grounds.

Nvidia Extends Its Chokepoint Into Its Customers' P&L

A revenue share on customer cloud income turns a depreciating hardware sale into a perpetual claim — and the same week, the market started charging Nvidia for capex headlines.

The structural item nobody indexed

The Information reports that Nvidia will take a cut of some customers' cloud revenues, which sounds like a footnote and is actually the business changing underneath the footnote: from selling an asset that depreciates over roughly six years to holding a royalty on the downstream revenue that asset generates. If it standardises, every GPU-cloud model built on buy-once-monetise-for-six-years is wrong at the gross-margin line, and terminal value resets across neoclouds, colocation platforms and GPU-collateralised lending.

Nobody has published the contract language. That is the diligence gap worth closing with portfolio CFOs directly, rather than inferring it from a headline that was written to be read quickly.


The reaction function flipped in the same cycle

The Information Briefing has the tape. Nvidia fell 5% on Monday on a $500 billion SK partnership headline and reporting about financing OpenAI's Ohio campus, and neither item was substantively new — the SK announcement largely reprised two early-June announcements, the Ohio financing was reported in early June. The one genuinely new fact, the Korea AI cloud doubling from 1GW to 2GW, was bullish and arrived inside a bearish tape.

Six months ago this headline adds five percent. In this cycle it cost five percent. The public comp set that late-stage AI infrastructure marks lean on has already repriced; the marks have not.

Price the instrument, not the number

Nvidia and SK have signed letters of intent only, and Nvidia declined to say which party contributes how much of the $500 billion, which is the interesting part rather than the number. The base rate is not theoretical: Nvidia and OpenAI signed an LOI in September 2025 for a 'landmark strategic partnership' and abandoned it within months. Any pipeline narrative resting on a non-binding agreement with Nvidia, a hyperscaler or a sovereign entity is carrying an option priced as a commitment.

The confirming cash number

Turing Post supplies the constraint that makes the rest legible. Alphabet posted its first-ever negative quarterly free cash flow of -$5.9B, even with Google Cloud revenue up 82% to $24.8B, against $195–205B of 2026 capex. AMD, meanwhile, is investing up to $5B in Anthropic as Anthropic commits to up to 2GW of MI450, and landed Anthropic and Azure in the same week. Vendors financing their own customers is a late-cycle demand-inflation marker; two hyperscaler wins for the second supplier is a genuine break in the single-accelerator narrative.

The honest counter-read: vendor financing may simply be cheap insurance on genuinely scarce supply, and AMD's two wins may stay at two. The claim that holds either way is narrower — cash conversion gets priced before any of this resolves.

What to do

  1. Pull actual contract language from every compute-exposed portfolio CFO within 30 days and stress-test margins at a 5–15% revenue-share overlay.

  2. Run an LOI audit across the portfolio and pipeline this quarter, converting every non-binding agreement with a chip vendor, hyperscaler or sovereign into a probability-weighted line item.

The bottom line

Stop underwriting model access and start underwriting who holds the model-selection decision — then re-price revenue quality at the application layer before your next mark.