Investment & Market Intelligence

The Investor

The Signal

Devin crossed $1B ARR the week its flagship benchmark was shown over 65% leaked.

When researchers rewrote SWE-bench problems so the code behaved identically but looked unfamiliar, agent pass rates fell up to 14.4 points and token burn rose 2.5x. The revenue behind coding-agent marks is real; the capability and cost-to-serve claims underneath them have never been cleanly measured.

In Play

  1. Anthropic's Opus 5.5 Leads a Three-Layer Price Cut

    Anthropic says Claude Opus 5.5 runs typical workloads about 40% cheaper than Opus 5, at $4/$20 per million input/output tokens, per TheSequence and Simplifying AI. The same week, Xiaomi released MiMo-V2.6-Pro free under an MIT license as the top open-weight model. Tencent priced Hy Image 3.5 at about $0.024 per image. The inference-heavy companies you back just got cheaper inputs, but only those with pricing power will keep the savings.

    Ask Clarity
    Try
  2. Coding Agents Hit $1B ARR on a Leaky Benchmark

    Cognition says Devin passed $1B in annualized revenue less than two years after launch, per TheSequence. Lovable grew from about $500M to $600M ARR between June and September, after raising its August Series C at $13.3B. TheSequence also cites research finding leakage in more than 65% of SWE-bench Verified, the benchmark coding-agent pitches lean on. The revenue in this category is real, but the capability and cost claims behind your coding-agent marks are not reliably measured.

    Ask Clarity
    Try
  3. Checking AI Output Is Becoming the Scarce, Priced Step

    Intel ended cash rewards on a decade-old bug bounty that paid up to $100,000 per report, per Risky.Biz, as big tech says AI-generated submissions are swamping triage teams. An r/sre thread flagged by Lex Neva questions whether automated root-cause analysis works without a human. Architecture Notes says agents now change code faster than people can review it. Value is moving from producing AI output to checking it. Intel has not explained its decision.

    Ask Clarity
    Try
  4. Agent Plumbing Turns Into Free Standards

    HarnessRouter, a free Apache-2.0 project, runs Codex, Claude Code, Hermes and DeepSeek Harness behind one OpenAI Responses-compatible API, per Simplifying AI. Architecture Notes reports that Google's ax orchestrator defines agent sandboxes with Kubernetes-style manifests, though its Agent Substrate runtime is not production-ready. Agent-infrastructure companies that sell basic routing or sandboxing now face free substitutes. Their defensible economics lie in idle-agent density, meaning how many paused agents one server can keep ready cheaply.

    Ask Clarity
    Try

Deep Dives

Opus 5.5's Discount Goes to Whoever Customers Can't Swap Out

A cut this deep only improves margins at companies with pricing power, and open-source routing tools have made that power harder for most AI apps to claim.

The 40% is two levers, and only one is guaranteed

Simplifying AI splits Anthropic's number into two parts: a 20% cut in per-token price and roughly 25% fewer tokens per task. Together they bring cost to about 0.6 of the old level. The price cut applies to every call. The token saving depends on the work being done, and Anthropic measured it on “typical” workloads. A portfolio company running long agent loops on unusual code may capture less. Three claims come from Anthropic's own measurements, not independent evaluations: the cost figures, the more-than-30% speed gain, and the claim that Opus 5.5 matches Claude Fable 5.1 on most tasks.

If you have exposure to the model layer itself, the arithmetic runs the other way. Simplifying AI calculates that each workload moving from Opus 5 needs about 1.67x more volume to keep Anthropic's revenue flat. Anthropic is pricing a Fable-class model below its own flagship. That is deliberate self-cannibalization aimed at winning share in coding and agents.


Why most of the saving won't stay with the companies you back

Every rival of a Claude-based app got the same cheaper model on the same day. It went live on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry. TheSequence's read is that the saving stays with an application only where the company has pricing power. Otherwise competition hands it to customers. Simplifying AI expects that pass-through within a few quarters.

The same week also lowered the cost of switching. HarnessRouter puts Claude Code behind one interface alongside Codex and DeepSeek's harness, so changing an agent's backend becomes a configuration edit. CopilotKit released OpenMuse, a free, MIT-licensed clone of Meta's new Muse agent. Architecture Notes reports that Google's ax is trying to set the convention for how agent sandboxes are defined. With switching this cheap, a company's claim to keep the saving has to rest on its workflow data, contracts and distribution, not its model access.

ReleasePrice moveBasis for the claimLikely beneficiary
Claude Opus 5.5 (Anthropic)~40% lower net cost than Opus 5; $4/$20 per million tokensAnthropic's own measurementsApps with pricing power
MiMo-V2.6-Pro (Xiaomi)Free weights, MIT licenseThird-party index: 46.32, top open-weight scoreManaged hosts able to serve >1T parameters
Hy Image 3.5 (Tencent)~$0.024 per image vs ~$0.12 implied for Google's Nano Banana ProTencent's own designersImage-heavy apps, if quality holds

The open-weight price floor needs one correction: free weights are not free inference. MiMo uses about 42B parameters per token, which keeps compute per token manageable. But holding more than 1T total parameters in memory takes multi-GPU serving that most enterprises won't run themselves, so that demand flows to managed inference hosts. Simplifying AI flags a possible limit, which it labels as its own inference: Chinese-origin weights may face procurement friction in regulated sectors. The source also does not report how MiMo's score compares with closed frontier models.

For image-heavy apps, Simplifying AI's illustration is concrete. At 10M images a month, an app would pay about $240K on Hy Image 3.5 versus about $1.2M at Nano Banana Pro's implied price. That saving only materializes if Tencent's self-reported quality holds up under independent testing.


The smart move

Treat the cut as a pricing decision each portfolio company must make on purpose, not a margin gain to book automatically. A company that plans to keep the saving should be able to say why its customers cannot route around it. A company that plans to pass it through should show the volume or share it expects in return. Either answer is defensible. Without one, a competitor makes the decision for it.

What to do

  1. Ask every Claude-dependent portfolio company this week to re-run inference COGS on Opus 5.5 using its own traffic, and to bring an explicit keep-or-pass-through pricing decision to its next board meeting.

  2. Commission a revenue-sensitivity update this quarter for any model-layer exposure, assuming workloads migrating from Opus 5 need about 1.67x the volume to keep revenue flat.

  3. Map managed-inference hosts that can serve 1T-parameter open-weight models in US and EU regions before your next AI-infrastructure IC, including each host's exposure to procurement friction over Chinese-origin weights.

Coding Agents Have Proven Revenue and a Broken Scorecard

Buyers clearly pay for coding agents, but nobody has yet measured what it costs to serve a customer's own codebase, and that cost decides the margin.

What rewriting the benchmark revealed

Researchers at SJTU, XJTU and ECNU built SchrodingerRepo by rewriting SWE-bench problems so the code behaved identically but looked unfamiliar. Pass@1, the share of problems solved on the first attempt, fell by 6.0 to 14.4 points across GPT-5.1, GPT-5.4-mini, DeepSeek-v4-Flash and Gemini-3.1-Flash-Lite. The agents also consumed more than 2.5x the input tokens, and 81.6–83.6% of their extra actions went to exploring the code. TheSequence's point is simple: every customer's private codebase is unfamiliar by definition. Decks built on benchmark scores therefore overstate capability and understate cost in exactly the setting that pays.

The unfamiliarity tax can swallow the model discount

TheSequence offers a rough illustration. Paying 2.5x the tokens at 0.6x the price still leaves cost per task at about 1.5x the naive estimate. The inputs come from different papers and setups, so treat the result as directional. For the coding-agent companies you back, the headline price cut does not settle gross margin. Two things settle it: how much an agent must explore before it can change a customer's code, and how much of that exploration the company has engineered away.

Some levers already exist. Salesforce AI Research's JIT Mem, a just-in-time memory method that fetches stored context only when the agent needs it, cut input tokens by 50.3–56.3% in testing. Researchers at Alibaba and Georgia Tech post-trained a mid-size open model, Qwen3.6-35B-A3B, with VHD-Play. Its agentic score rose from 0.204 to 0.815. TheSequence notes that JIT Mem's gain comes from its architecture, so memory alone is easy to copy. A company that owns neither lever pays the full tax.


Pricing the revenue proof

Demand is not in question. Cognition says Devin passed $1B ARR less than two years after general availability. Lovable added roughly $33M of net new ARR per month between June and September, about 20% growth in a quarter. Its $13.3B August valuation against ~$600M ARR works out to roughly 22x current run rate, on about 2x annualized growth. TheSequence proposes that multiple as the reference ceiling for the category.

CompanyScaleValuation markerOpen question
Cognition (Devin)>$1B ARR, under 2 years after launchNot disclosedGross margin on unfamiliar customer code
Lovable~$600M ARR, up from ~$500M in June$13.3B August Series C; ~22x current run rateWhether ~20% quarterly growth holds; users vs contracts

Two cautions apply. First, Lovable's claim to reach about two-thirds of the Fortune 500 is self-reported, and it counts individual users at those companies, not enterprise contracts. Second, Devin's milestone says nothing yet about gross margin, which is where the unfamiliarity tax shows up.


The smart move

Replace benchmark evidence with production evidence in every coding-agent process. Ask for three numbers on customers' private repos:

  • the merged-PR rate, meaning the share of the agent's proposed code changes that customers accept;
  • the review-survival rate;
  • the cost per accepted task.

If a company insists on benchmarks, ask for contamination-resistant ones built the SchrodingerRepo way. Any growth round priced above the Lovable ceiling should be tested on one question: does the company's growth rate match Lovable's?

What to do

  1. Replace SWE-bench Verified in your coding-agent diligence template now with three measures on customers' private repos: merged-PR rate, review-survival rate and cost per accepted task.

  2. Stress-test any growth-stage coding-agent round this quarter against Lovable's ~22x run-rate multiple, and require growth near Lovable's ~20% quarterly pace before accepting a higher price.

  3. Ask each coding-agent portfolio company this quarter how much of its exploration cost on customer repos it recovers through memory, post-trained models or other token-saving methods.

Cheap AI Output Is Making the Human Check the Scarce Product

Security, operations and software teams hit the same wall, which shows which layer's revenue grows with AI adoption instead of shrinking under it.

Security: one big payer just priced a finding at zero

Intel ran its bounty program for more than a decade, raised rewards sharply after the Spectre and Meltdown chip flaws, and paid academics studying side-channel attacks tens of thousands of dollars by their own accounting. Its page on Intigriti, the platform that hosts it, now reads “No bounty.” The last archived snapshot showing payouts is dated September 13. Intel has refused to explain the change. Catalin Cimpanu of Risky Business calls it the first domino. Companies that ran bounties for reputation now have a cost excuse to stop paying.

The chain here is simple enough to be suspicious of: AI multiplies findings, so each one is worth less. Pricing power moves downstream to validation, prioritization and remediation. Platforms that earn fees on payouts get squeezed, and their escape route is to become the AI triage layer. That is a different company than the one they are now. Risky Business names the second-order risk too: smaller legitimate payouts push top researchers toward exploit brokers, so more undisclosed flaws end up used in attacks. One payer is one payer. A category re-rating needs two.


Operations and code: the human check is still the product

In reliability engineering, an r/sre thread flagged by Lex Neva asked whether automated root-cause analysis (RCA) can “actually nail root cause without a person doing the final synthesis, or is that still mostly aspirational marketing from vendors.” It drew a large, engaged response. Plenty of AI SRE valuations assume the tool replaces that final synthesis outright. If buyers experience it as a copilot instead, the pricing model, expansion path and TAM shrink together. One thread is a hypothesis to test in reference calls, not a survey.

Architecture Notes describes the same wall in software teams. Agents now change systems faster than humans can review every diff, and nobody has solved how to help people follow the architectural decisions behind those changes. Ayman Nadeem built Nuanced around persistent plans. He has since concluded publicly that long AI-generated specs added friction without improving understanding. Founders rarely say that part out loud.


Judgment can be taught, so decision data becomes the moat

TheSequence's coverage of Taste-Bench sharpens the picture. Taste-Bench poses 502 questions about choices at decision points, and the best frontier model, GPT-5.6 Sol, scored 59.7%, with larger reasoning budgets buying nothing at all. A distilled Qwen3.6-27B advisor, meanwhile, lifted a fixed executor on SWE-bench Pro from 14.6% to 33.7%. That lift rests on only 41 held-out tasks. This is probably too clean, but the thesis follows anyway: if judgment can be learned from recorded decisions, then proprietary trajectory data, meaning logs of which choices worked, becomes a real advantage for whoever owns it. The counter-thesis is that 41 tasks is a rounding error and the advisor is an artifact of one benchmark's quirks.

DomainWhat AI made cheapWhat stayed scarceStrength of evidence
SecurityFinding bugsValidating and fixing themOne payer's unexplained cut, plus big-tech complaints
OperationsAutomated incident analysisFinal root-cause synthesisOne high-engagement practitioner thread
SoftwareWriting code and specsReviewing and understanding changesPractitioner and founder accounts
AgentsRaw reasoningJudgment at decision pointsBenchmark with a small held-out sample

The smart move

The diligence question worth stealing from Risky Business is whether a company's revenue grows with the volume of AI output its customers must check, or with the volume of output it produces itself. The first kind of revenue rides AI adoption. The second is being commoditized by it. Sourcing follows from there, favouring companies that pair the check with proprietary exploitation or decision data, which also means hours not spent on vendors whose entire pitch is raw throughput. Before re-rating bounty platforms as a category, wait for a second major payer to follow Intel.

What to do

  1. Add an autonomy audit to every AI SRE, AIOps and automated-RCA deal this quarter. Request the share of root causes accepted without human rewrite and the change in time-to-root-cause, and take reference calls with on-call engineers, not only the buyer.

  2. Request within 30 days, from any crowdsourced-security platform in your portfolio or pipeline, the revenue split between bounty-linked fees and subscription or triage fees, plus payout trends for its top 20 customers over four quarters.

  3. Build a sourcing map this quarter of validation and judgment tools whose revenue grows with the AI output customers must check, covering exploit validation, remediation automation, advisor models and agent-decision review.

The bottom line

Across security, operations, software and model pricing, the same economics showed up: producing AI output got cheap, while proving that output correct stayed expensive and human. That breaks the habit of reading every model price cut as margin for the companies you back, because the savings go to whoever customers trust to check the work, and everyone else passes them through. Rewrite your AI diligence template this week so each deal is scored on proof of correctness in the customer's own environment, such as accepted work and human rewrite rates, instead of benchmark scores, autonomy claims or a vendor's price sheet.