Product & Strategy

The Product Desk

The Signal

Chinese open-weight models now take ~60% of US companies' OpenRouter tokens.

DoorDash and Airbnb are running Kimi K3 and GLM-class models in production, not piloting them, per The Algorithmic Bridge. Treasury's Scott Bessent is floating sanctions on US adopters. The tradeoff is now explicit: the best available model against a dependency a policy move could strand overnight. Teams that know exactly which workflows sit on those models can price that risk. Teams that don't are guessing.

In Play

  1. Chinese Open Weights Are Already In Your Cost Base

    US companies now route nearly 60% of their OpenRouter token usage to Chinese open-weight models, per The Algorithmic Bridge, with DoorDash and Airbnb running Kimi K3 and GLM-class models in production. Kimi K3 delivers near-frontier quality at roughly a third of frontier API pricing. Treasury's Scott Bessent has floated sanctions on US companies that use Chinese AI. The same line item now holds your largest margin win and your newest policy risk.

    Ask Clarity
  2. Controllability Becomes a Product Requirement

    Reps. Lieu and Moran introduced the AI Kill Switch Act on July 23. It would require companies above $100M compute spend and $500M revenue to support government-ordered throttling, feature disabling, shutdown, and rollback, preserve model weights and telemetry, and report within 15 days, with fines up to $20M a day. Read it as a feature list, not a statute: enterprise buyers will ask for throttle, rollback, and audit logs long before enforcement exists.

    Ask Clarity
  3. Who Captures the Inference Savings

    An AI startup can book $100M in revenue and pass $90M straight to model providers, per TLDR Founders' reporting, and cheaper models do not fix it because customers either demand the savings or bring their own inference. Agentic features consume 10-100x the tokens of chat. One open-source tool cut token usage a median 82x by sending structural context instead of whole codebases. Efficiency is now a pricing decision, not only an infrastructure one.

    Ask Clarity
  4. AI Defaults Need an Escape Hatch and a Legal Review

    Google Photos shipped a one-tap toggle back to classic keyword search after users revolted against irrelevant Gemini results. The same week, OpenAI opened ChatGPT Health to every US adult across free, Go, Plus and Pro tiers, one day after a Florida pastor sued over advice to skip seeing a doctor. Techpresso reports 70% of health queries already happen outside the dedicated health hub, so the product boundary you draw does not constrain what users ask.

    Ask Clarity
  5. Machines Became the Majority Audience

    One development site logged 268,000 AI agent requests against 107,000 human pageviews over two months. ChatGPT-User drove 73% of that agent traffic and reads HTML almost exclusively, while Claude Code requested Markdown 76% of the time through an Accept header. The llms.txt file everyone recommends drew only about 37 fetches from named assistants. Serving each agent the format it asks for is cheap and measurable; the standard file is not.

    Ask Clarity

Deep Dives

Your Cheapest Model Is Now a Foreign-Policy Dependency

The sanctions case against Moonshot rests on a distillation timeline that does not hold, which is exactly why your written legal position has to exist before a buyer asks for it.

Start with what a founder actually did this quarter, not what the White House said. The accusation is that Moonshot AI distilled Anthropic's Fable to build Kimi K3. The Algorithmic Bridge lays out the calendar that argument cannot survive: Fable was available to users only June 9-12 and again from June 30, Kimi K3 shipped July 16, and K3 beats Fable on some benchmarks. A student model does not overtake its teacher in a two-week window. Anthropic separately settled its own copyright case for $1.5 billion. So treat the distillation charge as live policy risk, not as settled fact about provenance.

Why the usage is sticky. Separate the pitch from the thing being done. Chinese labs give the weights away and monetize hosted inference, so the model is free and the tokens are cheap. Newcomer reports Kimi K3 hitting near-frontier quality at roughly a third of frontier API pricing, with a 1M+ token context window, topping Moonshot's internal benchmark against everything except GPT 5.6 and Fable and leading on some coding and agentic tasks. TLDR Founders adds the adoption picture: open-weight models now sit in roughly 80% of startups and trail frontier capability by only 4-6 months.

Where the sources disagree, and why that matters. The Algorithmic Bridge puts Chinese open-weight models at nearly 60% of US companies' token usage on OpenRouter. TLDR Founders pegs open-weight models at 25-50% of volume across OpenRouter and Vercel. Both point the same direction, and the spread is the instruction: measure your own mix rather than paste either number into a planning doc.

The functional argument nobody priced in. During the Hugging Face compromise, OpenAI's and Anthropic's models refused to help because of safety guardrails, and Hugging Face ran GLM 5.2 on its own hardware to finish the forensics. For a product that touches security, operations, moderation, or incident response, refusal behavior is a requirement you test in a bake-off, not a philosophical position you admire.

CampPositionEffect on your roadmap
Treasury (Scott Bessent)Floats sanctions and pressure on US companies using Chinese AIAdoption carries documented policy risk
OpenAI policy (Dean Ball)Wants government to impose "large amounts of regulatory risk" and FUD on adoptersExpect vendor FUD inside competitive deals
CommerceWants to fund US open-source alternativesA domestic hedge may arrive, unevenly
Little Tech Alliance200+ startups including Y Combinator, Replit, Proton, Yelp lobbying against bansAn outright ban is not on the table today

The cost floor falls either way. Nvidia's Jensen Huang used his first-ever X post to back open weights, and Microsoft claims its in-house models cut costs up to 89% versus OpenAI. That is a vendor claim, not a benchmark. Whether the cheap tokens end up Chinese, American, or in-house, a margin model built on last quarter's frontier pricing is already wrong.

The move is architectural, not political. The decision this week is not whether Chinese weights belong in the stack. The decision is a 2x2 on control: can you switch inference providers inside one sprint, and do you know which features stall if a lab lands on an entity list. One investment covers both the price war and the sanctions order.

A model you cannot swap out in one sprint is not a vendor choice. It is a policy bet.

What to do

  1. Produce a one-page model-dependency map by end of week: every AI feature, its provider, its share of inference spend, and its US or China exposure.

  2. Run a bake-off against Kimi K3, DeepSeek V4, and GLM 5.2 on your top three real workloads this sprint, scoring quality, latency, cost per 1,000 tokens, and refusal behavior.

  3. Get a written legal and compliance position on Chinese-model usage before your next roadmap review, and name one owner to flag any sanctions or entity-list action within 24 hours.

Congress Just Wrote Your Agent Controllability Spec

Throttle, rollback, and forensic telemetry become procurement checkboxes whether the bill passes or not, and the connector defaults that make agents sellable are what make them dangerous.

A procurement lead reads the AI Kill Switch Act and does not see policy. She sees a requirements list. Reps. Lieu and Moran introduced it on July 23 as an amendment to the Homeland Security Act. Covered companies would have to support government-ordered throttling of inference and compute, feature disabling, full shutdown, and rollback to earlier versions, while preserving model weights and telemetry and reporting within 15 days. Penalties run up to $20M per day, DHS gets emergency authority, and the thresholds are $100M in compute spend and $500M in revenue. Passage is uncertain. The buyer expectation is not. Procurement will ask for these controls well before any enforcement date.

Two of those requirements are harder than they read. Forensic telemetry means more than what teams keep now. SANS research notes that the standard 48-96 hour retention windows already miss correlated events. Rollback means a versioned, known-good model state, which most AI features cannot produce. Prompt templates, model versions, and tool schemas drift independently, and nothing pins them together as one artifact.

The connector story is the nearer-term risk. Zenity Labs reported through Bugcrowd on June 4 that a single crafted ChatGPT link could feed instructions into OpenAI's agent builder, auto-wire the victim's existing connectors across Outlook, Teams, Slack, SharePoint and Google Drive, switch off approval prompts, and publish a scheduled agent. That agent read attacker emails tagged "TASK," searched files, took passwords and API keys, and sent phishing as the employee. OpenAI patched it in four days by removing the offending URL parameter. Automatic connector wiring and skip-the-approval onboarding both live on someone's backlog right now. Teams file them as activation improvements.

A second failure mode belongs in the same epic. Agent memory can be quietly poisoned so a corrupted belief lies dormant and triggers harmful action later, which defeats filters that only inspect immediate inputs. Agents that remember need memory integrity validation and audit logging of their own.

ControlWho is commoditizing itYour call
Prompt-injection scanningMicrosoft, natively in Defender for Office 365 before Copilot reads emailBuy — do not rebuild the filter
Agent runtime governanceOpenBox, one SDK across LangChain, LangGraph, n8n, MastraBuy if you are on those frameworks
Secret leakage at the promptSequirly browser extension scanning prompts and uploadsPoint solution, likely absorbed by platforms
Connector scoping and approval integrityNobody — still your design surfaceBuild; this is the differentiation window

Why buyers care now. MIT Technology Review reports AI-generated errors reaching an official court transcript, reportedly the first known case, with judges also taken in by AI misinformation. The failure mode enterprises fear has shifted from "the model was wrong" to "a human trusted it anyway." That changes what a demo shows. Controllability and a measurable human-review catch rate become the thing on screen, and the disclaimer becomes the footnote.

The controls worth having early are the throttle, the rollback, and the audit log; they show up on procurement questionnaires before they show up in law.

What to do

  1. Scope a controllability epic this quarter covering throttle and pause, versioned rollback to a known-good model state, and telemetry retention beyond the current 48-96 hour window.

  2. Audit connector permissions on every agentic surface this sprint: least-privilege defaults, explicit user consent before wiring Outlook, Slack, Teams, SharePoint or Drive, and tamper-resistant approval state.

  3. Kill or gate any setup-via-link agent configuration flow behind authenticated confirmation before your next agent release.

The $90M Question: Who Keeps the Savings When Inference Gets Cheap

Cutting your token bill is the easy half; the hard half is designing a price your customers cannot claw the savings back out of.

A founder books $100M in revenue and watches $90M of it route straight to Anthropic or OpenAI. TLDR Founders' number is the moment worth sitting with, because cheaper inference does not land in your margin by itself. Here is what teams tell themselves: when the model gets cheaper, the savings are ours. Here is what customers actually do: they demand the discount at renewal, or they bring inference they already pay for. The routing win then shows up on the customer's invoice, not your gross margin. Unless the price is anchored to something other than tokens.

The token curve is moving against you on exactly the features you are adding. Agentic harnesses — coding agents, autonomous multi-step workflows — consume 10-100x more tokens than conversational features. A COGS model built in the chat era is off by an order of magnitude for anything on the agent roadmap. Separate the thing being pitched from the thing being done: a market has already formed to arbitrage the gap. Spot and secondary venues now trade idle GPU capacity and unused credits at 20-80% discounts to retail inference pricing, which makes procurement a lever the finance partner can pull.

The efficiency lever that customers never see. An open-source tool, code-review-graph, achieved a median 82x reduction in token usage by feeding assistants precise structural context through Tree-sitter instead of re-reading whole projects, and it re-indexes a 2,900-file codebase in under two seconds. That gain never appears on the price list, which is precisely why it stays yours. The caveat is sharp. Prompt-caching savings are fragile. Adding a tool, reordering a schema, swapping a model, or rewriting conversation history silently erases them. Routine roadmap moves can destroy an inference budget without a single alert firing.

LeverWhat it changesWho captures the gain
Route simple calls to a cheaper modelCost per call fallsYour customer, if you price per token
Context efficiency and cachingSame output, fewer tokensYou — invisible to the buyer
Package the savings as a paid tierCost control becomes a featureYou — Cursor gated its router to Teams and Enterprise
Price on usage, seats, or outcomesRevenue decouples from token costYou, structurally

The pricing assumption under attack. Fugu-Ultra v1.1 shipped a capability upgrade across coding, agentic and reasoning work at the same price as v1.0. Users are being trained to expect smarter AI for free, so any tier priced on model quality is a depreciating asset. The counter-move is already on the board. Cursor kept its cost-savings router on Teams and Enterprise plans only, converting an efficiency gain into an upsell instead of a giveaway — and leaving individual and SMB developers unserved if you want that wedge. That is the tradeoff, named in the same breath as the play.

Buyers are formalizing the question. Cursor, Google Cloud and BMO convene at Harness's FinOps Excellence Summit on July 29 specifically to connect tokens to outcomes and push cost guardrails into deployment pipelines. When a bank shows up to demand cost-to-outcome proof, that proof becomes an RFP line. "We don't measure cost per outcome" becomes an answer you have to give in writing. The forcing function here is the RFP, not the summit.

If your price is a markup on tokens, every efficiency win you engineer belongs to your customer.

What to do

  1. Rebuild the cost model for every AI feature on the backlog this sprint as cost-per-active-user and cost-per-outcome, using 10-100x chat token volumes for anything agentic.

  2. Add an AI cost-regression check to the release gate so prompt-cache-breaking changes — new tools, model swaps, schema reorders — are flagged before they ship.

  3. Re-anchor at least one AI price point off model quality and onto usage, seats, or outcomes before your next pricing review this quarter.

Bots Outnumbered Humans 2.5-to-1, and They Read Different File Formats

The cheapest growth experiment available to you this quarter is content negotiation, and the discoverability tactic everyone recommends returns almost nothing.

Two agents crawled the same server for two months and asked for two different things. Over that window, evilmartians.com logged 268,000 AI agent requests against 107,000 human pageviews, and the agents did not behave as one bucket. ChatGPT-User drove 73% of agent traffic and requests HTML almost exclusively. Claude Code asked for Markdown 76% of the time through an Accept: text/markdown header. That is content negotiation — serving a different representation of the same page based on what the client asks for — behaving like a product decision with a measurable outcome, not a formatting detail.

The null result is the money-saver. llms.txt, the file every AI-discoverability post recommends, drew about 660 fetches, of which only roughly 37 came from named assistants. A hidden hint div got exactly zero attributable follows. If llms.txt is on the content-ops roadmap, that is capacity to reclaim today and spend on format matching instead. The caveat: this is one property over two months, so treat the ratio as directional and instrument your own logs first.

This is a distribution channel, not a curiosity. One in four software developers now build APIs for AI agents rather than for humans. Amazon adopted MCP — the Model Context Protocol, an emerging standard for how agents call external tools — for Alexa+ third-party integrations. Datadog built its Cloud SIEM agent tooling on MCP with progressive disclosure specifically to reduce context consumption, and Xweather ships an MCP-ready API with 15,000 free monthly calls. Products with structured, machine-navigable interfaces get reached by that ecosystem. Products without them are invisible to it.

SegmentShare of agent trafficWhat it wantsYour move
ChatGPT-User73%Clean semantic HTMLFix HTML structure on high-intent pages
Claude CodeCoding-focused, significantMarkdown via Accept headerServe Markdown by content negotiation
llms.txt fetchers~37 named-assistant fetchesNothing measurableDeprioritize entirely

The human door is narrowing as the machine door widens. Google's AI search increasingly answers the question before the click, favoring creator and question-based content over branded pages, and Reddit threads are rising as a cited source. On the demand side, 55% of social users post less and 47% have deleted apps from stress. Read alongside the agent numbers, the audience mix behind the traffic charts is shifting even when the totals look flat.

The order of work is the decision. Segment first, because you cannot A/B a format for an audience you have not measured. Then negotiate format, a days-long change with directly observable behavior behind it. Then restructure high-intent content around real user questions, where both the AI answer engines and the agents pull from. None of this requires a model, a vendor, or a budget line.

Who reads your docs is now a segmentation question, and one of your two largest segments does not render CSS.

What to do

  1. Stand up server-side segmentation of agent versus human traffic within one sprint, broken out by ChatGPT-User, Claude Code, and other named crawlers.

  2. A/B content negotiation on your top documentation and product pages this quarter — clean HTML for ChatGPT-User, Markdown for Accept: text/markdown — and measure citation and referral lift.

  3. Cut llms.txt work from the roadmap and reallocate that capacity to format matching and question-structured content.

The bottom line

Pick the single wrapper layer you can defend — provider swappability, provable cost per outcome, or user-facing control — and write it into one PRD as an acceptance criterion this sprint.