Product & Strategy

The Product Desk

The Signal

OpenAI's 80% price cut lands on a tier half of US business spend never touches.

Luna now runs $0.20 per million input tokens, which is a real number attached to a tier most production traffic never touches — Sol is where the calls actually go. The same week, Amazon raised hardware 60% and Nvidia signaled 15%+ server increases, so the in-house alternative you'd normally price against a vendor discount got more expensive at exactly the moment the discount became real. The forcing question for this quarter is not whether to take the cheaper tier. It's what share of your spend sits on the tier that didn't move.

In Play

  1. Tiering Beats Bargaining on Model Cost

    OpenAI cut its GPT-5.6 Luna tier 80%, to $0.20 per million input tokens and $1.20 output, and trimmed Terra 20% to $2/$12. OpenRouter's Peter Walker reports the frontier Sol model still absorbs more than 50% of US business spend inside OpenAI's family. For your gross margin, that gap is extraction and routing steps paying frontier prices. In the same week Amazon raised hardware prices 60% and Nvidia signaled AI server increases above 15%, so the move-inference-in-house thesis got weaker, not stronger.

    Ask Clarity
    Try
  2. Retrieval Is the Token Lever You Control

    Atlassian says customers pairing Teamwork Graph with Codex or Claude Code used nearly 50% fewer tokens building an app, and its Code Context release claims 44% more accurate agent results with 48% fewer tokens. Kiro credits an ~82% lower cost per completed task to its spec-driven workflow structure rather than to GPT-5.6 itself. Both numbers are vendor self-reported with no disclosed methodology. The implication for your roadmap is that the largest available cost cut sits in the retrieval and scaffolding layer you own, not in the model contract you negotiate.

    Ask Clarity
  3. Agent Telemetry Is the Unmodeled Line Item

    Three private data vendors posted growth curves that infrastructure software does not normally produce, and all three name the same cause: agents emit telemetry at machine scale. ClickHouse ARR reached $350M, up 40% since May; Cribl says it is near $400M, up 33% since February; Grafana is past $400M and reports half its customers used its AI assistant in 2026. Palo Alto Networks paid $3.35B for Chronosphere and Dynatrace $915M for Arize AI to own the layer that traces agent behavior. Your agent feature has a log-and-trace cost most PRDs never model.

    Ask Clarity
    Try
  4. Assistant Commerce Gets Its First Reported Lifts

    Walmart reported that Sparky users show 40% higher average order value, and Albertsons reported a 10% lift overall and 26% on complex queries. Benedict Evans flags the obvious caveat: both compare against a self-selected cohort, so the incremental number is unproven. Separately, an independent test of 20 verified Shopify UCP merchants found only 13 let an automated buyer reach checkout. You now have a fundable demand-side band and a broken fulfillment half in the same feature story.

    Ask Clarity
    Try
  5. Metering Rails Got Bought, Attribution Did Not

    Adyen paid $335M for Orb in June and Stripe paid $1B for Metronome before announcing OpenRouter at a reported figure between $7.5B and roughly $8bn. Brex CEO Pedro Franceschi told The Information that companies reselling tokens are probably not making much gross profit and may not know it, and confirmed Brex is not building a router. Build-versus-buy on usage billing has been settled by acquisition. Per-customer cost attribution has not, and it stays your problem whoever owns the rails.

    Ask Clarity
    Try

Deep Dives

The 80% Price Cut Nobody Is Routing To

Cheap model tiers collapsed and server hardware inflated in the same week, and teams without per-account compute attribution cannot act on either move.

The blocker is bookkeeping, not pricing

A finance lead exports the usage file, opens the AI line item, and cannot say which unit of revenue produced which dollar of compute. Brex CEO Pedro Franceschi gave The Information the sentence for that moment: companies "selling tokens and reselling tokens" are probably not making much money on a gross profit basis, and may not even know about it. The diagnosis is mechanical rather than strategic. Attribution between revenue and the compute that generated it has become hard to compute, which means the margin figure in most AI business cases is an estimate nobody can defend under questioning.

Workato's CIO Carter Busse names the operating cost of that gap: a standing Wednesday 2pm meeting with the CFO and head of engineering, held for no purpose other than reconciling AI invoices against usage dashboards that disagree with them. Separate the thing being pitched from the thing being done. The pitch is cost governance. The practice is three executives reading two documents that do not match. Busse also says the AI cost inside Workato's own agentic product "is going up quite a bit right now", which is a company with real cost discipline calling its own margin provisional.


The three tiers now have public evidence

TierPrice signalWhat belongs there
Frontier (GPT-5.6 Sol)50%+ of US business spend on OpenRouterNovel reasoning and high-stakes output only
Mid (Terra)$2 in / $12 out per million tokens, down 20%Mid-complexity reasoning currently parked on frontier
Cheap (Luna)$0.20 in / $1.20 out per million tokens, down 80%Classification, extraction, routing, tool selection
Open weight on reserved computeHugging Face past $150M annualized, up 50% in two monthsHighest-volume, lowest-stakes calls

The blue-chip reference is AT&T. Its CIO told the WSJ open models run 25% of workflows, and days later its VP of data science told The Information they handle 40% of employee queries, routed through LiteLLM specifically to curb Anthropic bills. Hugging Face's growth is coming from compute and storage rental rather than hub subscriptions, which is the receipt for volume moving off per-token frontier APIs. Lambda sells the shape openly now: frontier for the hardest 10%, open weights on reserved compute for the other 90%.


Where the sources disagree

Everyone agrees the leak is real. They disagree on the fix. Stripe and Ramp bet the router is the control point, and Ramp gives its router away free, which is a strong signal that routing itself is not a moat. Franceschi rejects the premise outright: "Routing to different models is actually not the problem... The problem is, how do you know the performance of the model that you're using is actually equivalent to that of a more expensive model?"

You cannot responsibly downgrade a step from the frontier tier to the cheap tier without a score to point at. The eval is the permission slip for the margin.

The second constraint is hardware. Amazon raised hardware prices 60% citing the memory shortage, Nvidia told customers to expect AI server increases above 15%, and Gartner projects DRAM revenue up 246.6%, with memory overtaking non-memory chip revenue in 2026. Tokens deflate while iron inflates. Any "bring inference in-house for margin" initiative got weaker twice this month, and Morning Brew's read of the tape has Micron down 5.83% on the Nvidia price news, which is the market pricing that cost transfer downstream to everyone shipping a per-user AI feature. The forcing function is small enough to run this week: list every model call in the product, and for each one write down the eval score that would justify moving it a tier cheaper. Calls without a score stay frontier by default, and that default now has a published price.

What to do

  1. Pull the last 30 days of model spend by call type and publish a three-tier routing map this week, naming which steps stay on the frontier model and which move down.

  2. Tag every model call with customer ID, feature and model before the next pricing review, then produce a gross-margin-by-account report.

  3. Re-run any self-hosted or on-prem inference business case at Amazon's +60% hardware and Nvidia's +15% server increases, and record an explicit go/no-go.

Retrieval Is Cheaper Than a Smaller Model

Two vendors now claim 40-50% token savings from what the agent is fed rather than which model reads it, and one has already warned its free context layer may start metering.

Why the mechanism is dull, and therefore generalizes

An engineer asks the agent which service owns a failing test, and the agent re-derives the link between the ticket and the repo from scratch. Then it does the same work again on the next question. A knowledge graph stores labeled entities and typed connections between them, which turns that relationship into a written-down fact instead of an inferred one. Less inference, fewer tokens. The part that decides whether this reproduces anywhere else is that the graph gets built automatically by connecting to business applications. Manual edge-drawing is what killed graph adoption for a decade. That is why the effect should show up outside Atlassian's marketing deck, and why a two-week spike is the right way to find out.

Separate the thing being pitched from the thing being measured. Atlassian's Code Context version of the claim is the testable one: search spanning GitHub or Bitbucket alongside Jira, Confluence and Loom, under existing repository permissions, producing 44% more accurate results with 48% fewer tokens. Kiro reports the same shape from a different layer, an ~82% lower cost per completed task, credited to its plan/build/test/review scaffolding rather than to the model underneath it. Two independent claims, both locating the savings above the model.


The pricing trap sitting under the free tier

The context layer is being monetized in two incompatible ways. Atlassian charges nothing for Teamwork Graph but has said billing may shift to how often the AI taps the database. ServiceNow, whose knowledge graph shipped in April 2026, already charges when customers point outside AI like Claude Code at its graph. Free-now-metered-later is the playbook. It lands in the COGS model as a variable nobody has priced.

The counter-position belongs in the bake-off. Databricks argues that lakehouse retrieval beats a standalone graph on completeness and real-time freshness, and it stores far more proprietary customer data than any app-layer vendor. The demand signal, meanwhile, is credible rather than narrative: Neo4j's Q2 2026 revenue exceeded its entire FY2025, with CEO Emil Eifrem noting that after twenty years of "screaming into the void about the value of graphs," everyone is suddenly talking about it.

A 25% token reduction from retrieval beats most roadmap items competing for the same engineering weeks — and it beats degrading the feature to protect margin.

The cost of concentrating the context

Consolidating repos, tickets, wikis and video behind one agent-queryable surface raises the blast radius of a single injection incident. PromptArmor built a malicious Copilot skill, a third-party-authored instruction set that is structurally a browser extension, and used it to pull data from Outlook, SharePoint and Teams to an attacker-accessible proxy server. Microsoft patched it and says customers need take no action. Its own purpose-built malicious-skill scanner returned a false negative. Claude, ChatGPT and Perplexity all support the same skill pattern.

So the forcing function for any roadmap carrying a plugin surface or customer-configurable instruction sets is simple: detection-based scanning is demonstrably not the control, while allowlisted egress destinations and human review before marketplace launch are. Both efficiency numbers are self-reported with no disclosed methodology. Treat them as a hypothesis worth two weeks of instrumentation, not a planning assumption to quote at an exec team.

What to do

  1. Run a two-week bake-off instrumenting tokens-per-resolved-task on your top agentic workflow, current retrieval against a graph-backed context layer, before the next model upgrade decision.

  2. Add a metered-context-access scenario to the AI COGS model and secure a 12-month price cap in writing at your next vendor renewal.

  3. Gate any agent skill or plugin marketplace launch on allowlisted egress destinations plus human review of third-party instruction sets.

Assistant Commerce Has Benchmarks; Agent Checkout Doesn't

Two retailers published order-value lifts you can take into a business case, while an independent test shows the automated-purchase half of the same story is not yet fundable.

Walmart published the spec as well as the number

A shopper asks Sparky for a weekly high-protein meal plan. Seconds later she has recipes and meal kits with one-click basket addition, and the assistant recognises the ingredients she already bought, online and in store, so she does not buy them twice. That last clause is the product. Cross-channel purchase memory is hard, the retailer owns it, and it belongs in v1. A generic chat surface bolted onto a catalogue reproduces none of it.

The reported lifts give a band rather than a target: Walmart at +40% AOV for Sparky users versus non-users, Albertsons at +10% overall and +26% on complex queries for its website chatbot. Both are quarterly-reported, which is what gets a business case approved. Both are almost certainly comparisons against a self-selected cohort, which is what should design the experiment. The smartest analyst in the review will find that caveat before the second slide. Ship with a randomized holdout and report incremental AOV. Anchor to the 10-40% band to get funded, then measure incrementally so the number is still alive in Q3.


The failure is on the fulfilment side

LayerEvidenceWhat it supports
Assistant-influenced basket+40% and +10%/+26% reported AOV liftsA funded v1 with a holdout attached
Agent-completed purchase13 of 20 verified UCP merchants reached checkoutNothing yet — treat badges as marketing claims
Acquisition mixOrganic sessions 140.1M to 125.4M across 54 brandsRestating targets in conversions, not sessions

An independent test of 20 Shopify UCP merchants carrying a "verified agent-ready" designation found only 13 allowed an automated buyer to reach checkout. The tester flags 65% as an optimistic ceiling, because a scripted browser is more reliable than a real AI agent. Self-certification is behaving the way self-certification always behaves. Separate the thing being pitched from the thing being done: the pitch is an agent completing the purchase, and the observed behaviour is a script reaching checkout roughly two times in three. A roadmap narrative that depends on the first is currently unsupported by the ecosystem underneath it.

The metric that hides the win

Brainlabs data across 54 client brands shows organic sessions falling from 140.1 million to 125.4 million, a 10.5% decline, after AI Overviews spread. Visitors arriving from AI tools converted at 1.5x the organic rate. A tenth of the traffic can disappear while revenue grows. A dashboard reporting sessions as the top-line acquisition metric will generate a false alarm and hide a real win in the same week, and the SEO-versus-AI-referral investment call cannot be made on blended data. Split the source and report conversion inside each one. Sessions become a diagnostic rather than a verdict.

The demand-side numbers are now good enough to fund the assistant. The certification badges are not good enough to fund the checkout.

What to do

  1. Rebuild the assistant business case on the 10-40% AOV band with a randomized holdout, and report incremental rather than cohort-comparison lift at the next review.

  2. Split AI-referred traffic out of organic in analytics and restate acquisition targets in qualified conversions before the next planning cycle.

  3. Gate any agent-checkout feature on a measured end-to-end completion SLO of 90% or better using a real agent, with mandatory human handoff on failure.

The bottom line

The pattern across these items is that every lever that actually moves an AI feature's margin sits outside the model: what you retrieve before the call, what you meter after it, and what you can prove per account when finance asks. That retires the planning assumption that vendor price cuts and hardware curves will repair unit economics on their own — the savings are real, but only teams with per-feature cost telemetry can claim them, and the same numbers also appear in enterprise finance and security reviews as purchase criteria. Name one owner this week for cost per completed task, broken out by feature and by account, and require every AI roadmap item to arrive with that figure attached before it earns engineering time.