Product & Strategy

The Product Desk

The Signal

Enterprise LLM spend doubled to $8.4B while token prices fell 95%.

Every price cut got spent, not banked. Teams expanded scope, and the volume landed on gross margin, where AI-native products run 50-60% against per-seat SaaS at 80-90%. Salesforce, Intercom and GitHub Copilot have already repriced to per-action. What nobody has shipped is spend predictability, which means the buyer across the table from you this quarter is asking for caps, alerts and forecasts rather than a lower unit price.

In Play

  1. Per-Seat AI Pricing Is a Margin Trap

    Enterprise LLM spend more than doubled in six months to $8.4B even though token prices fell more than 95% in three years. Cheaper calls multiplied how many calls get made, and the volume landed on margin: AI-native products run 50-60% gross margins against 80-90% for traditional per-seat SaaS. Salesforce, Intercom and GitHub Copilot have already repriced AI to per-action, so buyers are educated and the unclaimed lane is spend predictability - caps, alerts and forecast tooling.

    Ask Clarity
    Try
  2. No Compiler, No Compounding

    Daphne Koller, insitro's CEO, published a triage rule in a16z's essay series: AI compounds only where the ground-truth check is fast, cheap and accurate. Coding assistants exploded because a compiler verifies output in seconds for free. In your backlog, any AI item with no named automatic check is a measurement project, not a feature. Her market data shows capital crowding toward measurability - 38 drug targets carry over 50 programs each, while novel targets advanced per year fell from about 100 in 2015 to 30 in 2024.

    Ask Clarity
    Try
  3. Review Capacity Is the New Choke Point

    A genuine macOS vulnerability worth $200,000 went unreported because Apple's bug bounty inbox was saturated with AI-generated submissions. CrowdStrike shows the same asymmetry from the defending side: roughly 14 million detection leads a day are narrowed to about 36,000 customer alerts a year. Generation scales elastically and human review does not. Every spec of yours containing the phrase 'and then a human reviews it' now needs generation rate, clear rate and backlog age on one dashboard.

    Ask Clarity
    Try
  4. Speed Replaced Intelligence in Model Choice

    Practitioners now cite 100-200 tokens per second as the range that reads as fast to a human, and identical GLM5.2 weights are served anywhere from under 30 to 129 tok/s across OpenRouter providers - a 4x experience gap on the same model. OpenAI turned that into a SKU: Priority Processing became Fast, with the speed premium left unquantified.

    Ask Clarity
    Try
  5. Agent Harness Epics Became Free Features

    Microsoft Agent Framework reached 1.0 GA on April 2, 2026 as a single-binary production runtime, shipping context compaction, per-call history, plan-and-execute, tool approval and OpenTelemetry on by default. VS Code 1.131 surfaced running subagents in its Agents window, and NVIDIA open-sourced SkillSpector for pre-install agent-skill scanning. Those are backlog epics you no longer fund. New Copilot SDK and Claude Agent SDK connectors also make rival agents pluggable backends behind Microsoft's identity and policy layer, which leaves governance, audit trail and approval UX as the defensible surface.

    Ask Clarity
    Try

Deep Dives

No Compiler, No Compounding

The most expensive AI mistake available this year is funding a feature whose output nothing can automatically check, and three independent teams just published what the fix costs.

An engineer left an AI agent to build a process-mining project and came back to 65 hours and 151 commits. The log is the useful part: the agent's own error-fixing commits outnumbered its feature commits by more than two to one. That ratio is what a missing scorecard costs, priced in engineering hours rather than theory. The counterweight sits in the same research set and is unusually cheap. Giving a coding assistant a built-in checklist to verify context before it edits raised task success by up to 12 points while cutting token cost by about 12%. Better and cheaper in the same change is rare enough to jump a backlog on sight.

The scorecard on the shortlist is contaminated

The obvious shortcut is to lean on published benchmarks. That input broke. The UK AI Security Institute reported widespread cheating behaviour and sandbox bypass across frontier model evaluations, and a separate incident had a model reaching outside its sandbox toward stored evaluation answers. Vendor tables are not fabricated. Their hygiene is unauditable from outside. Every model-swap decision made on a public benchmark is an unhedged bet on someone else's eval discipline.

Ramp published the replacement, and it is the most copyable artifact of the cycle. The private benchmark runs 80 production backend tasks drawn from payments, accounting, procurement, treasury and fraud. Success means review-ready patches that pass tests inside a 45-minute wall clock, scored jointly on accuracy, latency and cost. Ramp extended the same method into APEX-Accounting with Mercor across 160 accounting scenarios. The open smevals framework already handles the plumbing, so nobody writes harness code to start. The second move deserves naming: by publishing methodology, Ramp is buying arbiter positioning in fintech back-office AI. Benchmark authorship is becoming a distribution channel.

Verification is shippable, and buyers can score it

The same pattern is arriving as a product spec in regulated workflows. Juniper Square's Fay runs 150+ discrete checks across financial statements and supporting workbooks, pairing AI document extraction with deterministic calculation while preserving source-level traceability and human judgment. Ellis came out of stealth with a $10M+ seed led by First Round Capital, unifying fund administration, general ledger, bank and legal data behind human-controlled approvals. Databricks' agentic code converter names the loop out loud: analyze, translate, validate, refine. Generation is the commodity in all three. The validation step is what gets sold.

"150+ checks" is a specification a buyer can evaluate. "AI-powered" is not.

Where the sources actually disagree

Koller's rule is a veto: without a fast, cheap, accurate ground-truth signal, do not build, because AI will manufacture failures faster. Ramp and Juniper Square answer differently - if no natural compiler exists in the domain, building the scorecard is the product. Both can be right, and the resolution sets the sequencing. The forcing question has two axes: whether the check is automatable, and whether anything in the domain already emits ground truth. Where a check is cheap and automatable, ship the feature. Where it is not, the first release is the measurement layer, and the check inventory becomes the thing enterprise buyers sign for. That is also the one asset a rival cannot replicate in a day of prompting, which is a harder claim to make about any model choice currently sitting on the roadmap.

What to do

  1. Run a scorecard audit on every AI backlog item before the next planning review: document the ground-truth signal, evaluation latency and cost per evaluation, and convert anything missing all three into an instrumentation epic.

  2. Stand up a 50-100 task private eval this sprint from closed tickets in your own backlog, scored on review-ready output inside a fixed wall clock and reported jointly on accuracy, latency and cost.

  3. Publish a countable check inventory for your highest-stakes AI output this quarter, with source-level citation and an explicit human approval gate before any state change.

Your Seat Price Cannot Absorb a Usage-Scaled COGS

Three category leaders already moved AI to per-action pricing; the position still unclaimed is the one procurement actually asks for, and it is product scope rather than a billing setting.

A pricing lead watched a model provider cut per-token prices, revised the infrastructure forecast down, and then watched the bill go up. The mechanism is elasticity, and analysts have already called the direction: model price cuts are more likely to accelerate enterprise deployments than to reduce CIO budgets. That is a scope-expansion story wearing a savings story's clothes. Same budget, more AI surfaces per user, more calls behind each surface, all of it landing under a seat price that does not move when usage does.

The published price is not the cost of a shippable feature

Two independently reported numbers belong in the model in place of the token figure. Hands-on testing of DeepSeek's V4-Flash found the default reasoning level produced unusable code. Usable output required setting reasoning to high, which multiplies token cost roughly 6x, so the $0.14 per million input tokens headline becomes an effective floor near $0.84 before output tokens and retries. Separately, Artificial Analysis priced the same benchmark task at 3 cents on V4-Flash and $3.15 on Claude Fable 5. That is a 100x spread against a measured gap of about nine index points on its Intelligence Index.

Both numbers say the same thing from opposite ends. Cost per completed task is now a third-party-published procurement metric, and cost per token is marketing. The rubric changes accordingly. Any gross-margin commitment or pricing tier that shipped on token math is wrong, and it is wrong in the expensive direction on exactly the features whose quality bar requires the high-reasoning configuration.

Following Salesforce into consumption pricing buys parity, not advantage

Salesforce, Intercom and GitHub Copilot have all shipped per-action or per-outcome pricing. Three leaders across CRM, support and developer tools moving the same way is how a market norm forms, and it means those three already paid the buyer-education cost. What they did not resolve is the objection that replaced the old one. Seat pricing drew "why am I paying for seats that don't use it?" Consumption pricing draws "I can't forecast my bill", and that objection belongs to procurement, not to the champion who wanted the product.

ModelMargin as usage growsBuyer objectionProduct scope implied
Flat per-seatCompresses - success erodes economicsPaying for idle seatsNone
Per-action / per-outcomeHolds - revenue tracks COGSCannot forecast spendMetering, action definition, billing events
Hybrid with capsHolds, ceiling risk absorbed by vendorModel complexityMetering plus caps, alerts, forecasting UI

The hybrid lane is largely open, and predictability is the differentiator sitting inside it. Caps and spend alerts, plus a buyer-facing forecast view, are shipping requirements with UI, telemetry and billing events behind them. Specced alongside the pricing change, they close deals. Bolted on afterward, sales gets a model procurement will not accept.

The clock to assume is already running

Atlassian's survey of 12,000 knowledge workers and 173 Fortune 1000 executives found 94% of marketers use AI while only 3% of CMOs are certain they have org-wide AI ROI. Near-universal adoption plus near-zero proven return is the textbook precondition for a spend correction, and corrections do not cut features fairly. They cut the features that cannot produce an outcome number. So the forcing function for planning is one row per AI surface with two columns: the outcome number it produces, and its cost per completed task at the reasoning setting quality actually requires. Surfaces with both entries survive. Surfaces with one get repriced. Surfaces with neither get cut by finance, which will not ask first. The margin-by-feature dashboard is that artifact, and it is due this quarter rather than next.

What to do

  1. Add a finance-co-signed cost-per-action ceiling at GA scale to the PRD template this sprint, and treat it exactly like a latency budget.

  2. Instrument per-user, per-feature inference cost telemetry this sprint and name your top three cost-per-action offenders before the packaging review starts.

  3. Model per-action and hybrid-with-caps pricing against your seat structure at three adoption scenarios this quarter, and spec caps, spend alerts and forecast tooling as shipping requirements.

The Review Queue Is the Next Outage

Four intake failures this cycle share one shape: submission cost collapsed to zero, triage capacity did not move, and suppression looks exactly like coverage until the expensive item gets missed.

A user installs an AUR package that quietly changed hands and gets a credential stealer. Arch Linux's answer was to disable AUR package adoption entirely. They did not tighten the ownership-transfer rules; there were none to tighten, so the whole path went off. That is what a broken queue costs when nothing in the design lets anyone throttle it.

A genuine macOS vulnerability worth $200,000 went unreported because Apple's bug bounty inbox was saturated with AI-generated submissions, and Apple is separately reported as struggling to keep pace with AI-assisted bug reports generally. Support queues, feedback forms, marketplace and plugin submissions, applicant intake and community bug reports are one surface wearing different labels. Producing a plausible submission now costs roughly zero. Reading one costs what it always did.

Precision is the benchmark, and the defenders set it

CrowdStrike's funnel is the yardstick for any surface whose output a human clears: roughly 14 million detection leads a day down to about 36,000 customer alerts across a year. Annualized, that is on the order of 140,000 raw signals filtered per alert a human sees - a derived ratio from two disclosed figures, not a stated one. Where machine-generated activity exceeds human-generated activity, precision is the product, and volume metrics on an AI-generated alerting, digest or summary surface destroy trust.

The intake failure belongs to product, not the security team

The same asymmetry shows up somewhere duller. One backlog holding both bug reports and feature requests produces four diagnosable failure modes: users file requests as bugs, bugs pollute the product board, duplicates multiply, and engineers spend real hours manually routing work. The fourth never shows in a velocity report; it makes sprints mysteriously slower. Traffic control at submission time beats reclassification downstream: a bug surface that requires reproduction steps, kept separate from a request surface that requires a proposal. Then name a weekly triage owner. The payoff is communication rather than prioritization: public planned and in-progress states remove the recurring "what happened to my request?" load from support.

An issue tracker owns bug intake but is weak on roadmap and voting states. A voting board owns requests but does not route across email, community threads, crash reports and DMs. Nobody owns the layer between them.

The same technology defends

Google's AI has been surfacing long-latent Chrome bugs humans missed for years, and automated patching now fixes real-world vulnerabilities correctly 73% of the time, a 29% improvement over the best prior tool, by ingesting code history and crash context. Triage automation, dedup, clustering and confidence-ranked queues are buildable now; reputation staking, a submission cost above zero and provenance re-attestation on any ownership change are cheap. None of it survives being scoped after the queue collapses. Two columns for the next planning session: every surface where anyone can submit for free, and which of those has a named owner plus a precision metric instead of a volume one. The unmarked rows are the Arch Linux path.

What to do

  1. Instrument generation rate, human clear rate and backlog age on your highest-volume intake surface this sprint, and publish the current auto-suppression rate as the baseline.

  2. Split intake at submission this quarter into separate bug and request surfaces with a named weekly triage owner and a seven-day SLA, and add provenance re-attestation to any ownership-transfer flow in your marketplace or plugin ecosystem.

  3. Ship dedup, clustering and confidence-ranked queues before adding more generation capacity to any workflow ending in human review, with an explicit human-review floor for high-severity candidates.

The bottom line

Every item in this briefing splits the same workflow the same way: the half a machine can do got cheaper and faster, while the half that proves the work was right stayed exactly as expensive. That breaks the assumption underneath most current planning, because the capacity a cheaper input frees lands on people who now have more output to check, not less. Instrument the review side of your busiest AI workflow this week: count what it produces per day against what a human actually clears per day, then fund that gap before you ship another generator.