Product & Strategy

The Product Desk

The Signal

USC and Arizona State steered model abstention by up to 52 points with no retraining.

Solvability and refusal turn out to be separate directions inside the model, mean cosine 0.087. Instruction tuning optimised refusal against harm labels, not against whether a question had an answer, so the two never got wired together. Until now, moving a model's "I don't know" rate meant filing a ticket and waiting for the vendor. It is now a config value you can A/B against task success this quarter.

In Play

  1. Abstention Becomes A Parameter You Can Ship

    Researchers at USC and Arizona State found that models linearly encode whether a problem is solvable, at 0.939 mean probe AUC across 11 models, per The Sequence's read of the paper. That direction is nearly unrelated to the safety-refusal direction (mean cosine ≈ 0.087), and steering it moves abstention by 33 to 52 percentage points. So "I don't know" stops being a model-upgrade request and becomes an inference-time control with a testable rate and a regression ceiling.

    Ask Clarity
    Try
  2. Salesforce Deletes Its Own Agent Brand

    Salesforce is removing "Agentforce" from some of its top product names less than two years after making it the centerpiece of sales and marketing — and doing it days before Dreamforce, The Information reports. Enterprise buyers have absorbed enough agent demos to stop rewarding the word. Every agent claim on your pricing page, in-product labels and sales deck now needs a number behind it: resolution rate, cycle-time reduction, or cost per unit of work.

    Ask Clarity
  3. Tool Calling Hits 12B-Model Parity

    Google Research published ToolGrad, and a Gemma-3 12B fine-tuned on its output scored 83.1 on the BFCL function-calling benchmark against Gemini 2.5-Pro's 83.2. Cognition's SWE-2 landed 64% below Fable 5.1 on cost by training every reasoning-effort level in a single reinforcement-learning run, so the whole cost curve moved rather than one operating point. The orchestration layer of your agent features is now a make-or-buy decision, not a frontier-API constant.

    Ask Clarity
    Try
  4. Capacity, Not Price, Gates The Roadmap

    Nvidia's filings show three customers each above 10% of sales drove 44% of revenue in the first half of its current fiscal year, up from two at 36% last year and none above the threshold in fiscal 2023, per The Information. Huang is investing in neoclouds to manufacture new buyers, which gives you a short window on committed-capacity pricing. The same reporting has SpaceX overhauling its data center build-out and personal AI app Instinct facing a compute crunch — availability, not token price, is what breaks a 2027 business case.

    Ask Clarity
    Try
  5. Filtering Lost To Verification In The Hiring Funnel

    Greenhouse data shows the average job opening drew 89 applications in Q1 2022 and 243 by Q1 2025, about a 173% rise, after one-click apply and generative résumés took submission cost to near zero, Morning Brew reports. Employers answered with filtering rather than matching: 90% of US employers now run AI screening that discards qualified candidates alongside spam. Read it as a controlled experiment on any submission funnel you own — screening is saturated and commoditized, while verification is scarce and therefore valuable.

    Ask Clarity
    Try

Deep Dives

Abstention Is A Parameter Now, And Reviewers Will Ask About Blast Radius

The research hands you a hallucination lever that needs no model upgrade; reliability practitioners just handed your enterprise reviewers the two objections most likely to stall it.

Why instruction tuning never wired this

An engineer files a ticket: the agent answered a question that had no answer. The model had the signal. A linear probe reading internal activations separates solvable from unsolvable problems at 0.939 mean AUC across 11 models, and that geometry is largely present before instruction tuning, per The Sequence's read of the USC and Arizona State work. Refusal was trained on harmful requests, so the unanswerable-question signal never got wired to it. The recognition direction and the safety-refusal direction sit at a mean cosine similarity of roughly 0.087, effectively unrelated.

Steering the recognition direction moves abstention by 33 to 52 percentage points. That turns a qualitative epic into two numbers a release can be held to: abstention rate on a labeled set of impossible or under-specified queries pulled from real logs, and a hard regression ceiling on answerable queries. Teams tell themselves this is a research quarter. The honest scope of the first test is one engineer and three days, and the output is a metric a reliability OKR can name.

Structure is substituting for model capability

The same pattern appears one layer up. Procedural Graphs, from Google with Georgia Tech and Peking University, store editable (procedure, relation, procedure) triples: a "what to do" graph where a knowledge graph stores "what is." They soft-bias a ReAct agent instead of hard-constraining it, and they self-evolve offline behind a validation gate that vets added, deleted and updated topology. PG is model-agnostic. It runs on whichever vendor a team already pays. A competitor can adopt it on the same contract this quarter, so the only advantage left is packaging speed. PG-guided Claude, Gemini and Grok solvers set or match best scores across HotpotQA, MultiChallenge, GDPval, ALFWorld, τ-bench, BFCL and EnterpriseArena, with the largest survival lifts on EnterpriseArena.

The two objections that decide the review

Reliability practitioners have already written the reviewers' arguments. Sylvain Kalache's comprehension debt holds that every routine incident automation resolves is a rep a human did not get, so competence erodes exactly where automation cannot help. Sai Sandeep Koneti's extension of blast radius from deployments to decisions is sharper: before a decision gets automated, define how far a wrong one propagates, and whether its output is absorbed back into the system to influence later decisions.

If the abstention is not logged with what the model saw, reviewers cannot audit it, and the wrong refusals will not surface until a customer finds one.

Both bodies of evidence converge on one point: the reliability work sits in the scaffolding around a model, not inside the next base model. Where they pull apart is what the PRD has to reconcile. The research optimizes for the model producing the correct refusal. The operators care whether a reviewer can see the input the system acted on and the reason it abstained. The forcing function: no abstention lever ships until the log carries both of those fields.

The attach surface nobody has claimed

Uptime Labs publishes on training early-career responders and Gremlin on Kubernetes disruption budgets. The training-under-automation product line is unowned. A team already holding incident data, system topology and resolution history can build a practice mode: replay a real past incident against a synthetic environment, let an engineer work it, then show what the agent did. Second product, same data, and it answers the strongest objection to the first. That is inference from a thin evidence base; five customer interviews will tell you whether the pain is funded or merely felt.

What to do

  1. Run a three-day abstention-steering spike this sprint on your highest-volume hallucination-complaint surface, targeting a ≥20pp abstention lift on a curated set of impossible and under-specified queries with under 3pp regression on answerable ones.

  2. Add a required Decision Blast Radius section to the PRD template this sprint: propagation scope, whether outputs feed later decisions, the confidence threshold for acting versus recommending, and the undo path.

  3. Benchmark procedural-graph scaffolding against a bare ReAct baseline on 50 of your own multi-step task traces this quarter before committing roadmap to it.

Salesforce Deleted The Word It Spent Two Years Selling

The rename is the visible half of the story; the invisible half is three labs privately drafting the audit rules that will decide what evidence an enterprise buyer accepts from you.

What buyers accept instead of the word

Pulling a load-bearing brand off top products immediately before your largest customer conference is not SKU hygiene. That is the timing you choose when the name has started costing you deals. Wherever you currently write "AI agent," a buyer now silently appends "prove it" — and only three numbers survive procurement: resolution rate, cycle-time reduction, and cost per unit of work. If a claim cannot carry one of those, cut the claim. For the next 60 days, while a category-defining competitor is mid-rename, displacement messaging built on deployed outcomes has an unusually clean run.

The political frame moved the same direction. The Information reports the AI conversation shifting from existential risk — bioterrorism, agent swarms, hacking — to tangible harms voters feel: job losses, erosion of critical thinking, and reduced parental supervision of children. Democrats are favored to take the House and possibly the Senate in elections a couple of months out, Obama has urged party leaders to focus on AI oversight, and Sanders has a bill that would jail people for building superhuman AI. That second list is not a lab problem. It is a spec problem: age assurance, guardian visibility, human-in-the-loop defaults, and designs that build skill instead of replacing it.

The rulebook is already being drafted, privately

Anthropic, OpenAI and Google have been in private discussions about jointly creating a testing and auditing standards body, The Information reports, and those talks predate Amodei's weekend essay calling for industry coordination — which makes the public advocacy an announcement layer on a process already running. At an OpenAI companywide town hall the same week, Altman told staff he supports such an organization and that the major labs will have to build it themselves, without U.S. government support.

Follow that to its delivery mechanism, because it decides your timeline. No federal funding means no federal enforcement, which means the standard arrives as voluntary conformance enforced by procurement: three new questions on a security questionnaire, roughly two to four quarters after any criteria go public, and a deal that sits until you can answer them with stored evidence rather than assurances.

Model providerReported postureYour exposure if the standard lands
OpenAIIn talks; Altman backs industry-funded termsLow — safe default for regulated buyers
AnthropicIn talks; publicly leading the agendaLow — strongest trust story in a deal cycle
GoogleIn talks; no public posture disclosedLow, but posture undefined
Meta / xAI / MistralNot reported in the coalitionModerate-high — standards written by rivals
Chinese labsNot reported in the coalitionHigh — likely excluded from enterprise checklists

Where the two readings disagree

One reading treats the pacing essay as a commercial non-event that changed nothing about the capability curve you plan against. The other shows the private process underneath it. Both are right, and the reconciliation is the actionable part: ignore the essay, build the artifact. The effort itself is fragile — early-stage, unannounced, no governance structure, no funding model, no membership list — and three direct competitors writing the rules invites antitrust and regulatory-capture scrutiny that could stall the whole thing.

The AI safety standard won't arrive as a law — it'll arrive as three new questions on a security questionnaire, and the team with reproducible eval logs answers them in a day instead of a quarter.

Which is exactly why you build the capability and not the certificate. Versioned evals catch regressions whether or not a coalition ever launches. Audit logs shorten incident response. A governance one-pager shortens enterprise sales cycles this quarter. None of that work is stranded if the talks collapse.

What to do

  1. Sweep the pricing page, in-product labels, PRD titles and sales deck for agent-as-differentiator language by the end of next week, replacing each claim with resolution rate, cycle-time reduction, or cost per unit of work.

  2. Stand up versioned eval provenance this quarter for every shipped AI feature: fixed test sets, pinned model and prompt versions, stored outputs, and timestamped results retained as per-release artifacts.

  3. Add a standards-coalition posture column to the model-vendor scorecard this sprint and quantify the ARR sitting on non-coalition models.

Google Published The Recipe For Undercutting Its Own Function-Calling SKU

Three published results put your inference bill under your own control this quarter, and one set of Nvidia footnotes says availability — not price — is what will actually break the plan.

The method matters more than the benchmark

ToolGrad's contribution is an inversion. The conventional pipeline invents a plausible user prompt, then sends search agents hunting for an API chain that satisfies it — often landing on a suboptimal chain with an unverified label. ToolGrad builds the chain, actually executes every call, then back-writes the prompt that would have produced it, so the label is correct by construction. Google reports a 99.8% pass rate across ToolBench's 16,000 APIs against the query-first baseline.

That makes ToolGrad an evals methodology wearing training-data clothing, and it is the cheaper half to steal. Run the same inversion against your own tool surface — build and execute chains, back-write prompts, wire the set into CI as a release gate — and you have the measuring instrument every subsequent model swap needs. The commercial subtext is hard to miss: a lab published the recipe for replacing its own premium function-calling SKU with a 12B open model, which tells you where it expects margin to survive, and it is not in tool orchestration.

Cost-per-capability is now the competitive axis

Cognition's SWE-2 sits just behind Fable 5.1 on raw capability at 64% less cost, achieved by training every reasoning-effort level in one reinforcement-learning run rather than tuning a single operating point. Two consequences for your spec. First, expose reasoning effort as an explicit parameter behind your agent features, and consider mapping it to pricing tiers, because the entire cost curve is now tunable. Second, provenance became a procurement field: SWE-2 is post-trained from Kimi K3, Moonshot's open-weight base, and Cognition's own evaluation puts K3 at 54.5% on a Simplified Chinese censorship test against SWE-2's 95.2%. The risk is mitigable — but only for teams who test it, and that benchmark is the vendor's own, so replicate it before it enters a PRD.

Someone already ran your retrieval cost experiment

Pinterest published the full recall-versus-footprint curve at billion scale on 100 million GraphSage embeddings — the vendor-neutral evidence a serving-cost business case usually lacks.

OptionIndex reductionRecallVerdict
Product QuantizationHNSW −74%, IVF −93%70-80%Cheapest, materially hurts relevance
Scalar QuantizationHNSW −59%, IVF −75%above 90%Default for most surfaces
SPANN (SSD)40%+ CPU saved vs HNSW at 5B scale−5% vs DiskANN3x DiskANN QPS at a third the latency

Per-use-case selection through online A/B tests produced 20-30% serving savings, and the winning pattern is hybrid precision: quantize the on-disk embedding store, keep centroids at full precision. If you own search, recs or RAG, own recall as a product metric with a floor above 90%, or the 93% footprint reduction wins the infrastructure argument and your relevance quietly loses twenty points.

The counterweight: supply, not price

All of that says your cost curve is yours to engineer. The supply curve is not. Nvidia's own disclosures, which Martin Peers and Qianer Liu at The Information built out of SEC footnotes, show three customers each above 10% of sales driving 44% of revenue in the first half of the fiscal year ending July — against two at 36% last year and none above the threshold in fiscal 2023. A third whale ramped while total concentration rose only eight percentage points, so the demand mix is churning, not merely growing. Huang is now funding neoclouds explicitly to manufacture new customers, so a well-capitalized set of providers has an active mandate to win your workload. Meanwhile SpaceX is overhauling its data center build-out with a possible slowdown, and personal AI app Instinct faces a compute crunch severe enough to potentially trigger a round driven by GPU access rather than product milestones.

Capital is abundant and models are abundant; capacity is not — so engineer your margin now and underwrite nothing on tokens getting cheaper.

What to do

  1. Fine-tune a 12B-class open model on your three highest-volume tool-use flows this quarter and compare cost per successful chain against your frontier API at equal task success.

  2. Re-run unit economics for every backlog AI feature at flat and +20% inference cost before the next planning review, re-scoping anything that only clears the margin bar at −30%.

  3. Add base-model provenance and a content-neutrality evaluation to the vendor questionnaire this sprint, and pre-answer both in your own model card.

The bottom line

Read the pattern as one: the capabilities you used to rent — reliability, cheap orchestration, credibility in front of a buyer — have landed on your own backlog, while the two things no team can manufacture, physical capacity and a buyer's belief, are what now price the roadmap. That inverts the planning habit of treating vendor releases as your ceiling and your own engineering as the floor. Take your most-demoed AI surface this week, give it one number that survives procurement and one stored artifact that proves the number, and make both reportable before a customer asks.