Product & Strategy

The Product Desk

The Signal

AT&T's 56%-for-2% trade is the first exchange rate a buyer can hold your product to.

The mechanism is unglamorous. A LiteLLM router sends the hard tasks to frontier models and everything else to open weights, which covers 40% of 45 billion daily tokens today and is targeted at 60-70% with OpenAI and Anthropic spend held flat. Any feature you price on model quality now competes against that ratio, refreshed quarterly by finance.

In Play

  1. Open-Weight Routing Gets A Published Price

    AT&T told The Information it will hold OpenAI and Anthropic spending flat while running 40% of its 45 billion daily tokens on open weights, with a stated target of 60-70%. After adopting the LiteLLM router, its AI coding costs fell 56% while measured quality fell 2%. That is the first at-scale exchange rate an enterprise buyer can hold your product to, and the same buyer caps what each developer may spend on premium AI tooling.

    Ask Clarity
    Try
  2. Model IP Became A Rentable Input

    Nvidia agreed to pay $6 billion for a non-exclusive license to Poolside's model technology plus $1 billion in equity at a $12 billion pre-money valuation, per a leaked investor letter obtained by Newcomer. Non-exclusive means AMD, Google or AWS can license the same technology next quarter. Fortune's Term Sheet reports investors marking 10-20% of portfolios as vulnerable to foundation-model encroachment. Any feature whose edge is model quality is a two-quarter asset, not a strategy.

    Ask Clarity
    Try
  3. The Harness Beats The Model

    A research team wrapped database ACID semantics — validate before commit, isolate failed attempts, keep durable state — around a multi-step agent and beat Anthropic's Claude Code by up to 10.6% on a data-agent benchmark, with no better model. Two separate studies found memory-based self-improving agents look materially worse once task order and evaluation variance are controlled. Reliability is an architecture decision your backend team can ship today, and 'our agent learns from every session' will not survive a technical buyer's diligence.

    Ask Clarity
    Try
  4. Agent Authorization Turns Into A Buyer Checklist

    1Password, Cloudflare and Google each published a competing agent-identity architecture in the same week, and Uber open-sourced its Agent Detection & Response system with a public benchmark covering 300+ tasks, 17 attack techniques and 133 MCP servers. Independent vendors converging on one problem in one week is a spec forming, not thought leadership. Specs that form in August arrive as procurement questions in Q1, so your agent PRD needs an authorization section before a security reviewer writes one for you.

    Ask Clarity
    Try
  5. AI ROI And AI Relief Split Apart

    SolarWinds found 84% of IT service management teams say AI met or exceeded ROI expectations, while 52% report their total workload increased anyway. An SSRN working paper on Chinese secondary students found generative AI raised homework scores 18% while closed-book exam scores fell 20%. Census pulse data shows 55% of US workers used AI on the job, and 13% of those users report no time savings or more work. 'Saves X minutes per task' is a weak renewal story when the saved time gets reabsorbed by validation.

    Ask Clarity
    Try

Deep Dives

AT&T Priced The Quality You Give Up For Cheaper Models

Routing stopped being a research spike: operator numbers, a flat-rate counterattack from Replit, and three published efficiency levers together reset what your margin defense has to contain.

What the 56% is actually measuring

A developer inside Ask AT&T opens the model picker, sees GitHub Copilot, Devin, Claude Code and Codex, and takes whichever one she used last week. She is not optimizing anything. Her spending is capped, so the choice that feels like a capability decision was already made as a budget decision. AT&T tiers by task complexity, not by user: frontier models generate code, cheap open models summarize code that was already written. Premium AI tooling ended up as a menu of interchangeable options competing under a ceiling someone is actively lowering. That lands on the pricing page, not the architecture diagram.

Mark Austin, the VP running AI for AT&T's 100,000 employees, told The Information the open-to-frontier capability gap runs six to 10 months and appears to be narrowing, and that AT&T is repatriating inference onto its own Nvidia and AMD hardware because it beats renting cloud capacity. That is a roadmap clock, not a benchmark note. A capability that needs a frontier model today plausibly runs economically on open weights in two to three quarters. Model-brand positioning is depreciating collateral.

The cost-adjusted frontier moved underneath the labs

Gemini 3.7 Flash posted 84.6% on ARC-AGI-2 at $0.25 per task and 95.5% on ARC-AGI-1 at $0.12. GLM-5.3 Max reached 1597 points in Code Arena WebDev at $3.65 per million tokens. The caveat travels with the numbers. Practitioners flagged the leading community coding eval as saturated, with all models bunched at the top, and Qwen3.8-27B regressed against its predecessor on offline factual recall while improving at tool use. Separate the thing being pitched from the thing being measured. A public leaderboard can no longer carry a model decision. A private golden set can.

The meter, not the price, is the objection

Replit folded usage costs into its flat monthly subscription, claiming subscribers can create up to 30 times more than before on OpenAI's GPT-5.6 Luna. At the other end of the same market, a $200-a-month plan is reportedly exhaustible in one heavy Codex day, with consumption continuing past the stated cap. Sources disagree here, and the disagreement is useful. One camp reads metered AI pricing as a competitive liability to remove. The other says reprice off token consumption rather than seats and define an overage mechanic. Both demands resolve to the same missing artifact: the P95 heavy user's true inference cost measured against a frozen quality baseline. Replit's headline capacity claim is also bound to one supplier's price sheet, which is gross margin outsourced.

Three levers that cost days, not quarters

  • Semantic caching benchmarked at a 57.1% hit rate, 55.7% fewer tokens and roughly 15 ms hit latency.
  • Gisting at about 40% lower end-to-end latency and 15% higher throughput, per Shopify's writeup.
  • Markdown tool output instead of JSON, at roughly half the tokens for the same payload.

One trap sits underneath all of it. Vendor prompt caching matches on an exact prefix, byte for byte: reorder two cached policy documents and both become misses, and the observed production pattern is a small number of blocks serving nearly all hits. A savings claim sourced from an aggregate cache-hit dashboard is a projection, not a result. The forcing function for this sprint is narrow. Take the two highest-volume AI calls, measure cost per successful task against the frozen golden set, and route to open weights anything that clears it. What fails becomes a frontier line item with a name attached.

Every AI feature you shipped without a routing layer is now provably about twice as expensive as it needs to be, and your enterprise buyer has the receipt.

What to do

  1. Instrument cost-per-successful-task plus a frozen quality baseline for every AI feature, broken out by feature, task and model, before your next pricing review

  2. Run a routing spike on your single lowest-complexity AI task within two weeks using a gateway plus an open-weight model, targeting a 40% cost cut at under 5% quality regression

  3. Model your top-decile user's real inference cost against your subscription price this month and publish the fair-use ceiling where a flat tier stays margin-positive

Nvidia Rented A Frontier Model, Non-Exclusively, For $6B

Nearly the whole team that built the weights walked to the buyer, investors are marking down a tenth of their portfolios, and a 30-person publisher just demonstrated what a durable moat actually costs.

The mechanics behind the price

Poolside had a six-week window at the end of 2025 to raise $2 billion for a 40,000 GB300 cluster. The round did not close and the cluster went away. Nvidia then hired 109 of the fewer than 115 engineers and researchers that CEO Eiso Kant said built the model. The founders stayed with the company. Their own words were "directionally correct in a race where capital requirements went vertical." The people went one way, the weights got licensed the other way, and the license is explicitly non-exclusive.

That word separates the thing being pitched from the thing being done. The largest strategic capital allocator in AI paid roughly 54% of a company's pre-money value for rights it does not get to monopolize. Discipline note: the terms rest on a single leaked investor letter reported by Newcomer with no on-record confirmation from either party. The strategic read is usable. The arithmetic stays out of a board deck until someone confirms it.

The rubric investors are already applying to product roadmaps

Fortune's Term Sheet supplies the scoring system. Kapital Ventures' Kamran Ansari names exactly two categories with any insulation left, regulated license businesses and businesses with proprietary, hard-to-access data, and adds that "absent those two things, everything else feels exposed." Underscore VC's Lily Lyman keeps a portfolio lane for companies "caught in the crosshairs by OpenAI and Anthropic," where the value prop is already obsolete. The cautionary comparable is Perplexity: "it was so molten-lava-hot. I don't think it's that special anymore because Google caught up extraordinarily fast." Vista Equity's Robert F. Smith said onstage that a small but real portion of his software companies "no longer have a right to exist."

Those are company-level judgments describing a feature-level mechanism. A differentiated AI feature dies by release note, and release notes ship on a vendor's cadence rather than a fiscal year.

What the moat looks like when somebody actually builds one

Every, a roughly 30-person publisher with no applied-ML team and no labeled-data program, collected 30,000 historical edits from its editor in chief, hill-climbed a prompt against them, back-tested the result on her past work, and shipped an @-mentionable agent that enters a live Google Doc and leaves suggestions. No fine-tune. CEO Dan Shipper had been trying since GPT-3 and the answer was "no, you can't" until three thresholds moved: instruction following, computer use good enough to operate inside a third-party surface, and headroom to back-test against a large historical dataset.

Two details make this the useful half of the story. Every doubled headcount from roughly 15 to 30 while automating everything it could, so the business case is expert leverage, not FTE reduction. And Shipper puts the half-life of a product built on top of frontier models at three to six months, with no answer beyond willingness to rebuild.

The cross-source pattern is unusually clean. An acquirer renting model IP as an input, investors pricing model-dependent value at zero, and an operator monetizing logged human judgment all land in the same place. The forcing question for a roadmap review this week has two axes: can a competitor buy this capability off a price sheet, and does the product accumulate a decision record that nobody else has. Features in the buyable-and-no-record cell have a three-to-six-month clock on them, same as Every's.

The moat was never the model partner. It is the tens of thousands of logged expert decisions most companies are currently overwriting.

What to do

  1. Label every P0 and P1 initiative this planning cycle as regulated, proprietary-data, embedded-workflow or model-capability-dependent, and read the last count aloud in the review

  2. Run an expert-decision-exhaust audit this sprint: inventory every system already logging 10,000+ expert judgments with before/after state, and rank by volume and whether one named expert dominates

  3. Push three clauses into your next model-vendor renewal: advance notice of material IP licensing to third parties, most-favored-nation pricing parity, and capacity and rate-limit guarantees

Three Vendors Just Wrote The Agent Permission Spec For You

A one-click assistant exfiltration hole took roughly eight months to patch, an autonomous red-team agent beat an AI code reviewer, and enterprise buyers now have benchmark numbers to quote back at you.

Where the three architectures actually disagree

They agree on the primitive: short-lived, scoped, logged grants that bind human plus agent plus intent. They split on where enforcement lives. 1Password keeps secrets away from the agent entirely; a trusted local app verifies OS code-signing and mints short-lived SPIFFE JWT-SVIDs from the Secure Enclave or TPM, so the agent holds a per-call grant rather than a credential. Cloudflare's Agent Access Model authorizes every action against the task's accumulated state, with a Trust Ratchet that irreversibly strips capabilities once protected data is touched. Google's Beyond Zero authorizes individual actions on resources through external policy decision points. Most SaaS vendors expose no such decision point, which is why it is blocked in practice.

That gap is the competitive story. Kane Narraway reads vendor adoption, not enterprise capability, as the bottleneck. Cloudflare concedes that multiplayer access control, one agent serving several principals with different permissions, is unsolved. Shipping MCP authorization hooks and a CAEP receiver for real-time revocation early is standards leverage, not hygiene.

Why the timeline compressed this month

Microsoft patched a critical one-click Copilot exfiltration chain, CoSnitch to its discoverer, roughly eight months after disclosure. It moved enterprise data with minimal user interaction and no obvious red flags. Separately, Wiz's autonomous Red Agent exploited a misconfigured GitHub Actions workflow in a Snowflake repository and left with Jira credentials, after AI security checks had already cleared the flaw. One false negative retires "AI review" as a sole security gate anywhere in a merge path.

The volume numbers are worse. A ransomware crew scaled from 10 to 100 attacks per week on AI parallelization. Dream Research Labs documented a July 2026 multi-agent campaign that compromised Asian government entities in four days across 12 waves with up to eight parallel sub-agents. Prompt injection is now treated as structurally unsolvable at the model layer, which is why Figma places safety controls at the tool layer rather than in prompt instructions.

The published business case is not a rip-and-replace

Figma's security team reported roughly 70% lower time-to-resolution on complex alerts and 20% fewer on-call pages from AI severity downgrading, across audit logs from AWS, Okta, GitHub, GCP, osquery and 100+ other sources, built on existing Panther SIEM and Snowflake investments. Two design choices are worth copying outright. Memory is tiered: case memory as a retrieval corpus of historical alerts, steering memory as behavioral guidance in markdown, procedural memory holding learned database schemas, which sharply cut schema-discovery queries. And agent-authored pull requests default to draft, one line of policy that buys an approval gate on every irreversible write.

Uber's ADR-Bench is the dry run that happens before a customer performs one: 300+ tasks, 17 agent attack techniques, 133 MCP servers, open source. Ant Group's free SingGuard-NSFA, 185 risk variants across seven domains, compresses pricing for anyone planning to sell "governed agent" as a premium SKU. The forcing function for this sprint is a grid. One axis: is the action reversible. Other axis: did a named human authorize it. Only the reversible-and-authorized cell runs unattended.

No security reviewer asks whether your agent is smart. They ask which human authorized the action, how long the grant lasted, and which log holds the record.

What to do

  1. Add a mandatory Agent Identity & Authorization section to your agent PRD template this sprint: no standing credentials on disk, task-scoped grants, per-action audit log, human-principal plus agent-actor attribution

  2. Audit already-shipped agent surfaces within two weeks for three defects: long-lived tokens held by the agent, irreversible writes with no draft state, and tool calls with no log entry

  3. Dry-run your agent against Uber's open ADR-Bench this quarter and ship an admin surface listing connected agents, data touched, per-grant logs and one-click revocation

The bottom line

The through-line across these sources: every layer a competitor can buy off a price sheet lost its power to differentiate, and the layers nobody can buy — logged expert judgment, verifiable stopping conditions, provable authorization — quietly became the only things a buyer will pay a premium for. That breaks the planning assumption that model selection is strategy. Model selection is procurement now, and someone in finance re-runs it every quarter with your invoice in hand. Spend this week labeling every AI item on your roadmap by what a well-funded competitor would need in order to copy it, then move everything whose honest answer is \"an API key\" to the bottom.