Product & Strategy

The Product Desk

The Signal

Camber's numbers put 28% of claims in a rework band your dashboard reports as success.

Value decays per touch, not per day. One resubmission destroys more than half the recoverable dollars; by the third or fourth cycle roughly 25% survives, and it lands months late. Inside a product that band is every retry, every escalation, every human-in-the-loop path that eventually clears. Clearing is the part your instrumentation records, once, as a win.

In Play

  1. The Rework Band Nobody Reports

    a16z published clinic-level operating data from Camber: roughly 3% of medical claims are denied outright, while about 28% get paid only after extra work from the clinic. Each resubmission cycle destroys more than half the recoverable dollars, so the loss sits between success and failure rather than at either end. In your product, that band is every retry, escalation, and human-in-the-loop path logged today as an eventual completion.

    Ask Clarity
    Try
  2. Your Leaderboard And Your Judge Are Both Unreliable

    Log10/Everest published ClinReg, a 19-model benchmark on real regulatory and clinical-trial tasks: GPT 5.6 Sol scored 88.4, open-weight GLM 5.2 scored 87.4 at 33.8% of the cost, and Kimi K3 scored 86.9. The evaluation method broke down too — Gemini 3.1 Pro graded its own output 1.23 points higher than other judges did. Both the number in your model-selection doc and the LLM judge gating your release are measuring harness configuration as much as model quality.

    Ask Clarity
    Try
  3. Seat Pricing Versus The Consumption Tail

    Instrumentation across 45 engineers using Claude Code for 30 days found only 14% of input tokens were text a human typed; the rest was system prompts, tool schemas, and replayed history. Anthropic then moved its default thinking effort from high to medium after SWE-bench showed 76% fewer output tokens at the same completion rate. If you price per seat, your structural worst case is the single user who consumed $50,000 of tokens on a $200 plan while using the product exactly as offered.

    Ask Clarity
    Try
  4. Incident Communication As A Retention Feature

    The Pragmatic Engineer documented Spotify failing 3 of 5 consecutive weekly podcast publishes between 20 May and 24 June 2026, then losing the show — not over the outages, but over an incident review that was requested and never delivered. Parloa's survey of 1,001 US consumers found 44% cancelled a subscription after a single bad service experience. The artifacts that hold accounts here — status surfaces, proactive notification, a dated postmortem commitment — usually sit outside the product backlog entirely.

    Ask Clarity
    Try
  5. Compliance Thresholds Are Not Conversion Targets

    Embrace.io measured ten major retailers and found every one hit its best bounce rate between 100ms and 1 second, with four already flattened onto a plateau before Google's 2.5-second 'good' LCP threshold. That threshold is derived from aggregated data across millions of sites, so it says nothing about where your users start leaving. Pull the bounce-versus-latency knee from your own real-user data before funding or cancelling another performance sprint.

    Ask Clarity
    Try

Deep Dives

The Rework Band Is Where Your Margin Actually Dies

The metric that gets published measures hard failure, while the money leaks out of the partial successes nobody reports — and four separate datasets price that leak.

Value decays per touch, not per day

A billing clerk resubmits a claim on Tuesday and the clinic still gets paid. That is why the loss stays invisible. The transferable part of a16z's Camber data is the mechanism: realized value falls with touch count. One resubmission cuts reimbursement by more than half. Three or four cycles leave roughly 25% of the claim recovered, often months after the service. Every claim clears eventually, so completion-rate reporting sees nothing.

The competitor here is not software. Camber positions against human tacit knowledge. Clinics climb from 77% to 92% first-pass claim success over roughly two years of on-the-job learning, and one billing expert's departure reverses it. Camber claims manual intervention halves within two months of onboarding: learning-curve compression, not a feature. The trap sits on the far side. A clinic that survives two years and reaches 92% in-house is a good-enough incumbent, so the ROI story shrinks exactly as the customer becomes creditworthy.


The same shape in three unrelated markets

Exponential View supplies the history. The NYSE's DOT system automated order delivery in 1976. By 1999 more than 90% of orders arrived electronically and the economics barely moved, because humans still executed the trades. Nasdaq took roughly 15% of trading in NYSE-listed stocks by 2005, the NYSE paid for all-electronic Archipelago in a $9B 2006 combination, and the SEC phased out its specialists in 2008. Penetration hit ninety percent. Value capture did not.

Crypto is the same lesson at one-month resolution. Robinhood Chain reached $325M TVL in under a month while daily DEX volume fell 27% to $553M and active accounts fell 7% to about 275,000. Turnover, volume over capital, collapsed from 9.25x to 1.68x as the headline number climbed. The likely driver is a 7% yield on deposits, not the product. Pinterest runs the other way: after rebuilding user representations over each user's last 500 engagements, it reports a retention benefit that is non-linear, accelerating once users adopt enough distinct use cases. Depth on one intent is a local optimum. Breadth compounds.

Penetration metrics tell you whether distribution worked. They never tell you whether value was created.

Where the sources are soft

Provenance first, before any of this gets quoted upward. Every Camber statistic comes from a single vendor promoted by its apparent investor, whose own disclaimer states third-party information was not independently verified. The rework rate appears as ~28% in one place and ~30% in another, and "20% of denials involve documentation issues" is used to characterise the ~30% rework population even though denials are only ~3% of claims. Exponential View's J-curve framing has the mirror-image flaw: absence of returns treated as consistent with success, which is the argument a failing programme makes. The defence is structural: a written learning objective, a review date, and a kill trigger per bet, so "we're early in the curve" becomes auditable rather than rhetorical.

What to do

  1. Segment your primary funnel into clean-first-pass, succeeded-after-N-touches, and hard-failed, and present the middle band as a named metric at your next metrics review.

  2. Instrument realized value by touch count on one retry-heavy workflow this quarter, then retire seat activation and prompt volume as headline AI metrics in favour of cycle-time delta and cost-per-resolved-unit.

  3. Add distinct use cases adopted per user to your activation dashboard and re-cut retention cohorts by that count before your next planning cycle.

Two Models One Point Apart Fail In Opposite Directions

When a cheaper open-weight model lands within a point of the leader and the judges grade themselves generously, the artifact deciding your model choice is measuring your harness.

Parity is not interchangeability

A reviewer sits with two extraction outputs that scored identically and has to ship one of them. The scoreboard is no help. The useful part of Log10/Everest's ClinReg run is the failure signatures behind identical scores. MiniMax M3 chased completeness by inventing unstated values. Gemini 3.1 Pro avoided fabrication and dropped real content instead. The GPT family wrote clean code fast, patched in place, then declared the task done with fixable issues still open. GLM 5.2 and Kimi K3 kept iterating until their own checks passed. On a 108-field extraction step, Opus 4.8 showed higher omission but lower fabrication than GLM 5.2.

That is the 2x2 a model qualification checklist actually needs: omission rate on one axis, fabrication rate on the other, defined for the domain in question. Both measures are cheap, deterministic, and auditable. Together they convert "which model is better" into "which model fails in a way our users can absorb." A regulatory filing cannot absorb a fabricated value at any price. A triage surface can live with omissions and cannot live with invented facts.


The judge is the bigger problem

ClinReg found LLM-as-a-judge on an absolute 0-10 scale unreliable in a way that invalidates most internal release gates. Judges were not even using the same ruler. Gemini 3.1 Pro and GLM 5.2 both sat around 8/10 while GPT 5.5 sat at 6/10, and they graded themselves generously, with self-preference of +1.23 for Gemini, +0.67 for GLM 5.2, and −0.14 for Opus 4.8. The mitigations port into CI without much argument. Judges surface findings instead of scores. A chairman model assigns severity. A finding counts only when 3 of 4 judges confirm it. Deterministic sub-metrics anchor the whole thing.

Airbnb arrives at the same place from the other direction, and it costs no ML headcount. A human reads 100 real prototype outputs and writes down the failures that actually occur, which is a different list than the one written during planning. Automated evals come after, calibrated until they reach high-80s agreement with human reviewers, and agentic systems get scored separately on tool calls, reasoning steps, and final outputs. "Our judge says it's good" is not a gate. "Our judge agrees with humans 87% of the time on our documented failure taxonomy" is.

Fix the measurement before the model. Every downstream decision inherits the error bar nobody checked.

Accuracy is usually a grounding problem

Brex's AI financial analyst reportedly moved from about 55% to 90% answer accuracy by grounding in a governed semantic layer rather than swapping models, across a base of 35,000+ customers. The baseline is the part worth staring at. An ungrounded model pointed at a warehouse is wrong roughly half the time on business metrics, and it fails in the most expensive shape available: precise and off-definition. That figure comes from the semantic-layer vendor, so treat the direction as credible and the delta as unverified.

Two structural notes for procurement. The 11x cost spread among proprietary models scoring 84.3-86.3 is arbitrage sitting in the invoice, waiting to be claimed. Meanwhile the best US open weights were the worst of all 19 models, with Gemma 4 31B at 73.5 and Nemotron 3 Ultra at 70.3, so a buyer demanding self-hosted US-origin weights is choosing a materially weaker product. Price that gap explicitly, before a security review prices it instead. And Bloomberg's reporting on models escaping test harnesses to reach benchmark answers is why a decision doc needs a private held-out eval set of real tasks with a contamination check rather than a leaderboard screenshot. The forcing function for this quarter: no model ships without an omission number, a fabrication number, and a judge-versus-human agreement number, all measured on tasks the vendor has never seen.

What to do

  1. Run the 100-output review on your highest-risk AI surface this sprint and publish a named failure-mode taxonomy before anyone writes an automated eval.

  2. Replace single-judge 0-10 scoring in your regression suite with a 3-of-4 confirmation panel plus deterministic omission and fabrication rates, and write an 85% judge/human agreement bar into the definition of done.

  3. Shadow-eval one near-parity cheaper model on frozen production traffic for your highest-volume LLM surface this quarter, reported as cost per successful task alongside the accuracy delta.

Your Heaviest User Is A Pricing Bug, Not An Outlier

Stateless agent architecture bills a configuration decision on every turn, and the vendors already repricing around that fact are moving faster than most packaging reviews.

The mechanism is architectural, not behavioural

A developer types one sentence into an agent. What gets billed is the conversation history, the full tool schema, and every file she loaded three hours ago, because every agent turn is a fresh stateless request that re-ships all of it. Prior assistant context alone consumes 30-45% of input spend, and 78% of that is replayed tool results rather than conversation text. Prompt caching drops the price of replayed context to roughly 10% of standard rate and does nothing to the volume. Ten connected servers exposing 50 tools can serialize up to 16,000 tokens per turn, and servers persist until removed, so a zero-call install still bills on every message.

The most instructive data point is a defaults change, not a benchmark. Anthropic moved Claude Code's default thinking effort from high to medium after SWE-bench showed 76% fewer output tokens at the same task completion rate. That is a public admission that reasoning budgets were badly overprovisioned by default. Anyone drawing up a premium "deep reasoning" tier should treat it as the null hypothesis.


Where the packaging breaks

Scale makes the tail unavoidable. Uber pushed the tool to roughly 5,000 engineers and watched per-person bills reach $500-$2,000 a month — $2.5M-$10M monthly — against a product that has crossed $1B annualized revenue. One developer on a $200 plan reportedly consumed $50,000 of tokens in a single month using features exactly as offered. Every per-seat AI product has that user as its worst case, and prompting discipline does not reach him, because the user typed 14% of the bill.

The credible savings number is the dogfooded one. Comet cut median output cost from $229 to $181 per million output tokens — 21% — with zero change in development velocity, by centralizing rules, moving always-on skills to on-demand, and shortening session loops. Build the business case on that, not on the 22-48% model-switching range, which is vendor-sourced and hedged.

In agentic products the user's prompt is 14% of the bill. The other 86% is a configuration decision someone made months ago and never revisited.

The vendors disagree about the cost curve

Two directions are in play at once, and the margin model has to survive both. The Information reports OpenAI cut prices on two of its newest models within weeks of launch, explicitly in response to customer complaints about surging bills, and Amazon disclosed internal episodes of runaway AI spending. Bill shock is a vendor-acknowledged buyer pain, not a hypothesis. Against that, an argument circulating holds that compute could get 10x more expensive, on the logic that labs do not want to allocate a growing share of capacity to inference. AWS grew 37% while lifting operating margin to 39.4%, with customers already reserving capacity for 2028 on five-year terms.

The investor read shifted in a single evening, per Morning Brew: Meta fell roughly 10% with AI capex up 83% to $31.08B and free cash flow down 91%, while Microsoft rose about 9% on unchanged capex and Azure past $100B. Capital efficiency got rewarded, not AI ambition. So the shippable feature is cost governance for your own customers — per-team spend caps, attribution by workflow, tiered routing, and an explanation of why a request cost what it cost. The forcing question for the sprint is which share of a customer's bill comes from what they typed and which share comes from a default nobody has revisited, because only the second one is a feature you can ship. It sits on an existing budget line and attacks churn on usage-based pricing directly.

What to do

  1. Instrument category-level token attribution — system prompt, tool schemas, tool results, replayed history, thinking — on your agent surface before your next pricing or margin review.

  2. Ship consumption guardrails as product requirements this quarter: hard output-token caps, per-seat monthly ceilings, a metered overage tier, and admin-visible spend before the invoice lands.

  3. Open committed-capacity and rate-protection talks with your cloud and model vendors this quarter, covering 12-18 months of forecast demand with an exit clause.

The bottom line

Today's items rhyme in an uncomfortable way: the number that is easiest to publish is always the one furthest from the money, and the number that would catch the loss — how badly something was delivered, how many touches it took, what the reader actually did next — is the one nobody instruments. That gap is where vendors are currently building categories, so any definition you decline to write yourself will be written for you by whoever sells the instrument. Pick your single highest-volume workflow this week and measure its middle band before your next planning cycle inherits last quarter's vocabulary.