Investment & Market Intelligence

The Investor

The Signal

A $200B lease repackages Anthropic's compute risk as Apollo and Blackstone credit paper.

Google designs the chips and guarantees them, Broadcom guarantees them again, and the tenant leases them back, which is a great many guarantees for one set of silicon. The interesting part is not the size but the relabeling: compute risk comes out the other side as credit paper. This is probably the right reading, but the consequence is uncomfortable — an LP holding both venture AI and private credit funds carries that exposure twice under two labels, and only one of the two mandates thinks it owns it.

In Play

  1. Inference Cost Inversion

    DeepSeek has warned users to expect a 'significant' price increase and Alibaba is reportedly seeking revenue share, per Benedict Evans. Both sit on the price-performance frontier that anchored every 'inference gets 10x cheaper each year' underwriting model. Meanwhile Meta's Apache-2.0 30B Muse Glimmer runs agent loops on one consumer GPU, and Nous Research cut agent token use 48-66% by collapsing twelve browser tools into one. The deflation still arrives, free to every competitor, which makes it a baseline rather than a moat.

    Ask Clarity
    Try
  2. Incumbents Cap the Price of Last Year's Wedge

    GitHub moved Code Quality to general availability at $10 per active committer per month plus AI usage, deliberately outside its Advanced Security bundle. An Apache-2.0 challenger, PR-AF, claims the #2 slot of 42 reviewers on Martian's Code-Review-Bench at zero license cost, and Spotify open-sourced Xirp, the agent orchestrator 1,300 of its engineers already run. Demand got legitimized and the price got published in the same week. Any devtools or orchestration deal in your pipeline modeling $40-plus per seat is now arguing against a public anchor.

    Ask Clarity
    Try
  3. AI Capex Moved Into Private Credit

    The FT reports a $200bn structure in which Google designs and guarantees TPUs, Broadcom sells and guarantees again, an SPV funded by Apollo and Blackstone debt owns the hardware, and Anthropic leases it. Anthropic separately signed a $10bn six-year compute deal with a Norwegian bitcoin miner, and Amazon has permits for 7.65GW of self-generated Texas gas. Microsoft booked $24.1bn of revenue from OpenAI in the twelve months to June, over half its AI revenue, from a loss-making customer it partly funds. An LP holding both venture AI and private credit owns that risk twice.

    Ask Clarity
    Try
  4. Prop Shops Are the Unpriced Compute Buyer

    Optiver posted EUR4.5bn ($5.1bn) of trading income and EUR1.7bn ($1.95bn) of profit on roughly 2,200 employees in its 2025 results, about $886K of profit per head, per The Pragmatic Engineer. Its US CTO says ultra-low latency is no longer a moat in itself, and the firm now invests substantially more in better models than in lower latency. Serious prop shops are spending hundreds of millions each on research clusters, which is why NVIDIA, Groq and Cerebras are courting them. The firms will never take outside capital; the return sits with whoever sells to them.

    Ask Clarity
    Try

Deep Dives

The Margin Tailwind Went Free and the Input Went Scarce

Every agent company underwritten in the last 18 months assumed permanent token deflation, and both halves of that assumption have broken in opposite directions.

The term that matters is revenue share, not price

A per-token price increase is a cost problem, and cost problems have engineering answers: quantize the model, cache aggressively, shrink the tool surface, route the easy turns to something cheaper. A revenue-share demand is a different instrument entirely. Alibaba is reportedly weighing exactly that ask, per Benedict Evans, which converts an application company's gross margin from a number its engineers control into a number its supplier reopens at will. Anyone who underwrote a 70% steady-state margin behind a vendor holding revenue-share optionality underwrote a counterparty's forbearance and filed it as a cost curve.

Evans reads the market as a supply crunch in which labs can name their price. The counterweight he reports from the buy side is that large enterprises now presume dual-sourcing, because, in his phrase, you cannot rely on anyone right now. Demand-side behaviour commoditises the model layer in the same quarter that supply-side scarcity lets suppliers raise price. A poor place to own equity and an unreliable place to buy inputs, at the same time.


What the free tier now covers

Meta's Muse Glimmer is the specific fact that resets the COGS line: 30B parameters under Apache 2.0, running on a single 24GB RTX 4090 or a 32GB Apple Silicon Mac, sub-20GB at roughly 4-bit quantization for 0.2-1% accuracy loss, and about 233 tokens/second on an RTX 5090 using DFlash speculative decoding for roughly 3x throughput. It was trained around the agent loop itself, planning, tool calls, self-checking, failure recovery. Nous Research came at the same line from software, collapsing twelve Hermes browser tools into one and cutting token consumption 48-66% with no accuracy drop. Databricks has published its own token-cost optimisation methodology, which is the tell: cost-per-token engineering is table stakes, not differentiation.

LeverCost effectVerified?Hard limit
Local 30B open weightsVariable per-token cost becomes amortized hardwareLicense and hardware specs verifiable; capability claim has no published benchmarksWill not match frontier reasoning on hard turns
Tool-surface redesign48-66% fewer tokens per taskReported with no accuracy lossOne-time refactor any competent team can copy
Hosted frontier APIRepricing upward, revenue share in playPrice direction reported, magnitude undisclosedCapability ceiling is the reason you pay

The honest caveat: a 30B model does not replace a frontier model on genuinely hard problems. Most agent turns are not hard problems. They are tool calls, retries and the self-checks between them, which is what this class of model was tuned for, and also where the token volume in every agent forecast was supposed to come from.


Where the two readings diverge

One reading says suppliers gained pricing power. The other says harness efficiency cuts revenue per task even as agent volumes climb. Both hold at once, price per token up, tokens per task down. A model vendor whose top line rests on agentic usage growth now needs a token-efficiency discount applied to the per-task assumption. An application company loses predictability in the variable-cost line in either direction, and it is the unpredictability rather than the level that breaks a five-year model.

The mirror image is the trade almost nobody is pricing. This is probably too clean, but zero-egress, zero-per-token local agents under a permissive license look like a structural advantage in healthcare, legal, defense and EU data-residency accounts, where cloud-API competitors cannot follow on price or on data handling. Two ways that goes wrong before the license does: the frontier gap widens faster than the harness improves, or enterprises decide dual-sourcing means two clouds rather than one cloud and one laptop. The unhedged tail risk is license durability: a business whose entire cost structure assumes permanent unrestricted access to one open-weight lineage is carrying regulatory risk it has not named.

If a holding's margin case improves only because tokens got cheaper, you own a market condition, not a business.

What to do

  1. Commission a token-cost sensitivity across the ten largest AI-native holdings within 30 days, modeling gross margin at flat, +25% and +50% inference cost plus a supplier revenue-share scenario.

  2. Require documented dual-sourcing and a model-abstraction layer as a diligence condition this quarter for any company where inference exceeds 15% of COGS.

  3. Ask every AI portfolio CEO for a written local-inference and data-residency plan by month-end, then score which ones could sell into healthcare, legal, defense and EU-residency accounts.

GitHub Set the Ceiling, Apache 2.0 Set the Floor

AI code review got the strongest enterprise endorsement available and a published price cap inside the same 72 hours, which moves the money one layer up the stack.

The packaging decision is the TAM claim

The price is the least interesting term here. What to underwrite is where GitHub put the SKU: Code Quality shipped outside Advanced Security, which is a monetization judgment dressed as a packaging detail. Microsoft is asserting that maintainability carries standalone willingness-to-pay, separate from the security budget, and that is budget-line creation rather than a feature reshuffle. It also hands challengers a clean total-cost wedge, though Microsoft keeps the option to collapse that wedge by re-bundling the moment anyone gets traction. Both things are true at once, which is usually how these end.

The legitimacy half arrived in parallel, and it removes the last procurement objection. Linux 7.2-rc7 landed on August 9 with more than 400 fixes Torvalds credits directly to automated scanning across nearly every subsystem, including an eight-year-old use-after-free race in ptdump that Syzbot flagged in June, with Claude Opus 4.8 credited in tracing the root cause. Human triage and sign-off remain mandatory on every AI-assisted contribution. When objections surfaced in July, Torvalds told dissenters they could fork or leave. The most scrutinized codebase in software now runs machine-assisted review as ordinary process, under a policy that survived open dissent.


Three positions, one published price

PositionPriceReal moatStructural vulnerability
GitHub Code Quality$10/active committer/mo plus AI usagePull-request context, default-branch debt surfacing, threshold rulesetsIncremental spend for existing security customers; no self-hosting or data-residency story
PR-AF (Apache 2.0)$0 license, self-hosted, model-agnosticVerifies every finding against source and drops what it cannot proveVendor-supplied benchmark, no managed SLA, ops burden, #2 on the benchmark it chose to cite
Frontier-model DIYPer-token, uncappedCapability ceiling and the kernel referenceToken cost compounds across agent loops; hyperscaler intermediation dilutes the vendor relationship

Sources diverge on durability, and that disagreement is the whole decision. One reading says $10 is an introductory number that rises once the category stops being contested. The competing evidence is that incumbents keep giving away last year's wedge: Spotify's Xirp orchestrates across Claude, Gemini CLI and OpenAI Codex, was validated by 1,300 internal engineers, and is now free to anyone. Underwrite the floor, not the recovery. Forty-two reviewers on a single public benchmark is competitive density, not greenfield.


Where the money goes instead

Spotify's own precedent is the template, or rather the more useful version of it: Backstage became the standard and the money accrued one layer up, in audit logging, SSO, cost attribution and SLAs. The equivalent layer for agents is already evidenced as a gap rather than a preference. A survey of roughly 336 papers found GUI and computer-use agents systematically lack error recovery, safety checks and auditability, which are architectural gaps rather than model-capability gaps. Separately, hidden-activation correctness probes proved extraction-method-dependent and unusable as a substitute for running tests. The cheap shortcut to agent verification is closed.

The tooling wave has names attached already: Prefactor on real-time run scoring and drift, Cekura on voice and chat agent QA inside CI/CD, FetchSandbox on deterministic replay across 60-plus APIs. This is probably wrong in the particulars, but the one worth funding is whichever becomes the audit system of record for agent actions, because regulated buyers pay for provenance and do not pay for a dashboard. Assume a 12-24 month window before platforms bundle the rest.

Underneath sits the demand driver nobody markets: agents solve each assigned task correctly and never plan for future features, silently compounding technical debt that slows both engineers and subsequent agents. Every dollar of agent-generated code manufactures future demand for refactoring, and for the debt visibility that makes refactoring schedulable. Discipline note: PR-AF's 0.706 recall, its #2-of-42 placement and the 3x-findings-at-10x-lower-cost claims are vendor-supplied and unverified — annotate them before they anchor a valuation conversation.

In devtools the money just moved from shipping agent output to proving and cleaning it.

What to do

  1. Re-underwrite every AI code-review and orchestration deal in the pipeline against a $10-per-committer anchor and a zero-license floor before the next investment committee, and reprice anything whose model assumes $40-plus per seat.

  2. Add one question to every agent-infrastructure diligence this quarter: why does an enterprise buy this instead of deploying a free orchestrator that 1,300 engineers already validated in production?

  3. Open a sourcing sprint on agent audit-of-record and action provenance, targeting five first meetings in 90 days.

The Best Compute Buyer in the Market Will Never Take Your Money

The most ROI-transparent GPU buyer in enterprise has no procurement committee, and the layer it now spends on has migrated from latency to research compute.

The 25% rule

The heuristic travels further than the financials. Optiver runs 30-40% of its roughly 950 engineers on platform work, against 15-20% typical at large tech companies, per The Pragmatic Engineer's reporting. Take about 25% platform engineering as the dividing line: above it accounts build, below it they buy. That line is probably wrong at the edges, and it still sorts accounts better than revenue does. Optiver forked asyncpg, contributed a nanosecond-precision timestamp type to Postgres, built PG Feed on the Postgres write-ahead log specifically to avoid Kafka's extra disk I/O and latency, and shipped its own AI gateway and MCP hosting platform internally this year.

That last item closes a beachhead. Selling agent gateways or MCP infrastructure into elite financial firms means pitching accounts that already built the product, which sends that go-to-market down into tier-2 financial services and mid-market enterprise, where nobody staffs 300 platform engineers. Then again, Databricks won the entire data platform at a firm operating at nanosecond precision, alongside Kafka and Postgres. Where these buyers buy, they buy the incumbent lakehouse, which scopes challenger theses to research and backtest.


The layer the budget moved to

Alex Itkin, Optiver's CTO US, says what the industry used to keep indoors: ultra-low latency is no longer a moat in itself. The retreat system went from seconds to nanoseconds over a decade, and that decade is now the floor rather than the edge. Budget migrated one layer up, into research compute, where serious prop shops spend hundreds of millions each. Hence NVIDIA, Groq and Cerebras courting trading firms, Hudson River Trading on a GTC stage discussing Blackwell, and Jump among the first to deploy Vera Rubin.

The counter-thesis deserves a hearing: chip vendors court prop shops for the logo, and the revenue never scales past pilot. Nine-figure research budgets argue otherwise, as does a buyer profile enterprise software rarely sees. No external customers, so no procurement committee and no external deadlines. Own capital only, so budget authority sits with people who can attribute a GPU purchase to profit and loss within weeks. Partnership structure, so nothing waits on a board cycle.


Two second-order repricings

  1. Talent costs more than the models assume. AI labs now recruit out of prop shops, not just Big Tech, because the binding constraint at frontier labs is low-level systems and hardware talent rather than ML researchers. Kernel, HPC and FPGA hiring lines are underfunded against a market clearing at prop-shop levels and above.
  2. Agentic coding turns headcount into compute. AI coding tools have started raising builds per engineer per day, against bare-metal CI clusters that require production-grade capacity planning. GitHub Actions exposes no queue times or utilization, so Optiver built bespoke observability on webhooks. A product wedge sitting on a demand model, and the CI category is not priced on it yet.

Two things belong in the documents before anything ships. Knight Capital nearly went bankrupt on a single bug that produced a $440M loss, so anything touching the live execution path carries reputational and legal exposure far beyond its ACV, which is one more argument for the research, backtest and CI paths where budget is growing anyway. And 2025's record profit reflects a favourable volatility environment; profitable strategies decay quickly. Prop-shop revenue belongs at under a quarter of any portfolio company's forecast: a proof point for adjacent high-performance verticals, not the terminal market.

The dozen firms that can fund nine-figure research clusters are customers, not targets, and they have no procurement committee to slow you down.

What to do

  1. Direct compute, inference-hardware and research-infra holdings to stand up a quant-finance vertical go-to-market motion this quarter against a named list (Optiver, Jane Street, Jump, DRW, Hudson River Trading, Citadel Securities, Two Sigma), leading with cost-per-research-cluster and ROI-per-GPU-hour.

  2. Re-run engineering compensation assumptions in every active infrastructure model this quarter, adding 25-40% to systems, kernel, HPC and FPGA lines.

  3. Commission five reference calls into fintech and HFT platform teams within 30 days to test whether agentic coding is measurably increasing build volume per engineer.

The bottom line

The items here rhyme in one uncomfortable way: the improvements the market underwrote as future margin are being handed out for free, while the inputs everyone assumed were abundant now come with a supplier holding pricing power and a contract date. That inverts the usual reading of a cost curve. A falling one no longer accrues to the company that needed it, and a rising one lands entirely on whoever sells per seat or per task. Stress-test every AI-application thesis in the book with zero cost tailwind and a supplier that reprices at will, because the holdings that still clear are the only ones whose growth was ever doing the work.