Product & Strategy

The Product Desk

The Signal

GitHub Copilot's metering took Pegasystems from $20K a month to $260K.

Annualized, that lands near $3.1M against a $10–15M total software budget. One line item taking a fifth of the wallet, with no contractual warning anywhere in the agreement. AI vendors took 8% of enterprise software spend this year, up from 1.4%. The executive absorbing that shock is the same one who approves the renewal on your metered feature, and he will be reading the usage curve rather than the pitch that justified it.

In Play

  1. AI Spend Is Funded By Cutting Your Line Item

    Procurement platform Zip reported that AI vendors took 8% of enterprise software spend in the 12 months to August 2026, up from 1.4% the prior year, while total software budgets rose a median of 13%. That leaves non-AI software spend growing roughly 5% — the AI budget is substitution, not expansion. Your next business case has to name the incumbent line item you displace and the executive who owns it.

    Ask Clarity
    Try
  2. A Decision Tier Appeared Below The LLM API

    TypeSafe removed the waitlist for Jev on September 21 and gave every new account $5 of credit that it says covers about 120M tokens — an implied ~$0.042 per million, against $2 per million for Grok 4.7's input. Jev returns typed choices plus probabilities in one forward pass instead of prose. That reprices routing, relevance scoring, moderation and injection screening — the per-item features most teams killed on unit cost. Pointer notes TypeSafe has published no benchmarks, pricing or customer list.

    Ask Clarity
    Try
  3. Shipped, Enabled, And Invisible

    James Stanier recounted in The Engineering Manager a feature that was deployed and enabled for 100% of customers with near-zero adoption, because the sidebar holding its entry point did not slide out by default. Every release-level indicator read green throughout. If your definition of done stops at 'enabled', your status reports are accurate and wrong at the same time — and nobody catches it until someone walks the journey by hand.

    Ask Clarity
    Try
  4. Model Vendors Are Fixing Their Costs Before You Fix Your Price

    The Information reports Anthropic is in early, non-binding talks to lease up to 1 gigawatt of capacity directly from Apollo-owned Stream Data Centers, filled with Broadcom/Google TPUs rather than Nvidia GPUs. Augment reports the listing moved from October to November so investors see Q3 results, against reported annualized revenue above $100B (third-party sourced and adjusted). Vendors carrying fixed lease obligations sell committed capacity rather than tokens, which opens a negotiating window that closes once sites energize.

    Ask Clarity
    Try

Deep Dives

Four Cents A Million Tokens Reopens Your Cost-Per-Item Backlog

A hosted model, seven open-source clones and two free trial windows expiring within ten days hand you a buy-then-abstract decision to make while the pricing that justifies it still exists.

The clone wave, not the price, is the durable fact

Within the same news cycle as Jev's launch, at least seven open-source projects built on or against it surfaced — Kev, reflex, DocJev, System One Harness, Laya-CoreML, pgbot and webctl — and two of them advertise TypeSafe-compatible APIs. Compatible clones appearing this fast make the hosted endpoint a commodity interface rather than a moat. Your portability cost is near zero while those compatible clones exist, and rises afterwards. Unwind AI derived the ~$0.042 per million from the launch credit ($5 covering roughly 120M tokens); model your business case at list price, because that number is promotional.

OptionWhere it runsProfileLock-in risk
Jev (hosted)Vendor cloud~$0.042/M implied from the promoHigh — DocJev, webctl and System One Harness need a TypeSafe key
Kev / reflexSelf-hosted GPU or Apple Silicon0.6B/4B/8B checkpoints on Qwen 3.5; frozen-model, no task fine-tuningLow — compatible API
Laya-CoreMLOn-device, Apple Neural Engine~5ms on an M3 Max; fastest bundle caps at 96 tokensNone — offline, zero marginal cost

That ~5ms figure is the interesting one: it puts a model decision inside a synchronous request path. The 96-token ceiling on the fastest bundle is a real design constraint, not a footnote — you fit the decision prompt to it or you accept the slower bundle.

What the interface shape actually changes

Jev does not write. Given context and a question with predefined answers, it returns typed choices plus calibrated probabilities in a single forward pass. Ben's Bites captured the positioning with Aaron Francis's line that did 51.7K views: ask your mom if the sky is blue and she answers instantly; ask your dad why and he talks for ten minutes. Every place your product currently pays generative rates for a yes/no, a category or a score — support routing, relevance re-ranking, auto-tagging, moderation, ad detection, injection screening, gating an agent's next permitted action — is a constrained decision wearing a generative cost structure.

Note what builders actually did with it. The highest-engagement work retrofitted primitives users already touch constantly: Cmd+F, copy-paste, drag-and-drop, Gmail search by intent, sponsor-skipping, and a real-time negative-comment filter that drew 383K views. Nobody shipped a new app. Retrofits need zero user education and move funnel metrics instead of "AI feature adoption" — which also makes them the cheapest way to prove the new cost tier internally.

Two do-not-builds and two clocks

The most popular integration is the wrong one. Wiring a decision model into coding-agent context compaction forfeits prompt-caching savings, and Ben's Bites reports Codex's native compaction is clean enough that any custom edge decays within weeks — cost goes up, advantage expires. Orchestration and memory lost their claim as differentiators in the same cycle: Google open-sourced AX, a declarative agent runtime built for billions of agent tasks per cluster, and Supermemory's teardown of a viral assistant's memory claims the pattern is reproducible in ~60 lines of git-tracked Markdown (treat that as surface probing, not confirmed architecture).

Both clocks are short. Lovable's Gateway routes to Jev free until September 27, 23:59 UTC, with no announced price afterwards. Cognition's SWE-2 is free across Devin Cloud Agents, CLI and Desktop until October 8.

Where the sources disagree

Pointer, covering TypeSafe's exit from two years of stealth, notes it disclosed no benchmarks, no pricing and no customers while claiming both speed and price leadership — an invitation to be benchmarked publicly. Ben's Bites adds that Epoch audited 15 benchmarks and found flaws in nine. The gate is therefore unchanged regardless of how good the tier looks: a 50–200 example golden set drawn from your own production data, and no vendor number in an acceptance criterion.

Portability on a commodity interface is cheap while the compatible clones are available and expensive afterwards.

What to do

  1. Run a constrained-decision audit this sprint: list every call site where a generative model returns a yes/no, a category or a score, with call volume, per-call cost and p95 latency, then re-cost the top three at list-price decision-model rates.

  2. Decide on Lovable's free Jev routing by September 27, 23:59 UTC, and ship anything you keep behind one internal decision interface with a self-hosted or on-device fallback configured.

  3. Cancel any custom context-compaction or bespoke agent-memory ticket this sprint and record the reason in the backlog so it is not re-proposed next quarter.

Your Rate Card Stopped Predicting Your Bill

Published per-token prices fell hard through the year while enterprise AI invoices multiplied, and the gap between the two is where your packaging either wins accounts or loses them at renewal.

The invoice that produced a public critic

A CIO opened a monthly coding-assistant bill and read $260,000 where the previous month had said $20,000. Nothing in the contract had changed. Annualized, that is $3.1M against a total software budget of $10–15M: one tool consuming 20–30% of an entire enterprise software wallet, arriving with no contractual warning shot. Pegasystems CIO David Vidoni went on the record that AI providers "put so much burden on companies, people like us, to detangle…what is the right way to use these" cost-effectively. He is now imposing employee-level spending caps. He built the rate limiter by hand because no vendor shipped one. Per-user caps and pre-invoice alerting are shippable features, and so is the attribution that tells him which team spent the money.

Microsoft's answer to the balk was 1 million credits, worth $10,000, for one month of Copilot Cowork, reportedly modest compared with what other customers received. Vidoni now plans to expand Cowork testing into finance and marketing. A placation credit turned into a land-and-expand motion. An agent roadmap scoped to engineering is pointed at the wrong department, because the budget moved into finance and marketing on a $10,000 credit.

Flat price, higher invoice

xAI shipped Grok 4.7 at Grok 4.6's exact price and speed across the API, Cursor and Grok Build, free inside Grok Build, with a no-cost switch for existing users. A team reads a flat rate card as flat costs, but Grok 4.7 was trained on problems that take many hours, so it verifies its own output before returning it and manages a 500K-token context. Simplifying AI argues the meter that matters is token consumption, not the per-token price. Capability gains came in uneven: +5.9 points on CursorBench 4.0 coding against +11 points on EEBench electrical engineering. That spread is wide enough that a general leaderboard says nothing useful about a specific domain. Ben's Bites ran the same release by hand and came back blunt and negative.

The Information Briefing supplies the aggregate: token prices dropped 75% while actual AI bills tripled, driven by agentic consumption running at a different order of magnitude than chat. Two questions worth running against a roadmap this week. Does the invoice move when agent usage moves, and can the customer see it move before the invoice arrives. Pega answered both with employee-level caps it wrote itself.

MoveBuyer responseWhat it teaches your roadmap
Shift a coding assistant to usage-based metering13x bill, CIO installs caps, accepts creditsMetering survives only where switching costs are high — and buys you a public critic
Discounts and free AI trials alongside metered AI (Workday, HubSpot, Amazon)Adoption rises on credits, not proven valueFree-then-meter is deferred churn dressed as traction
Hold list price, ship a model that reasons longerCost per completed job rises invisiblyCost-per-token is a vanity metric; cost-per-completed-task is the real one

Discount the sample, keep the direction

Applied AI is explicit that the 8% wallet figure comes from dozens of tech-forward companies averaging 2,400 employees: Snowflake, Datadog, Cloudflare, AMD, and OpenAI, which is simultaneously a customer of the platform and a named beneficiary of the shift. Treat it as a leading indicator for high-software-intensity segments that reaches mainstream enterprise two to four quarters later, not as current market-wide penetration. The corroboration for the direction is separate and independent. Computerworld reports incumbent software vendors are now discounting AI modules defensively to stop defection to frontier labs.

Flat pricing wins deals against metered incumbents right now. It stops winning them the month those incumbents smooth their own invoices.

What to do

  1. Ship cost-per-completed-task (tokens consumed ÷ successful outcomes, p50/p99) onto every AI feature dashboard this sprint and gate any model swap behind it.

  2. Model P95 consumption per account this sprint and flag every customer where an AI overage could exceed 20% of their total spend with you.

  3. Rewrite the budget-source section of your next business case this quarter to name the incumbent line item you displace and the executive who owns it, and promote per-team spend caps, usage dashboards and pre-invoice alerts to P0 on the admin surface.

Green Dashboards, Invisible Features

Release status, benchmark accuracy and volume-shaped success metrics all report supply; three separate threads show what happens when nothing in the stack measures the customer's outcome.

The sentence to steal

Shipped was true. However, findable by users was not.

The fix is a specification change, not a QA process. Release telemetry and journey telemetry are different systems, and "enabled for 100% of customers" is a deployment fact reported as a product outcome. A five-rung definition of done closes the gap: deployed, enabled, findable in the default state, first use within 14 days, week-4 retention — each rung with a named owner and a named telemetry source. Most teams stop at rung two and report rung five.

This gets structurally worse in 2026, because the layers that used to catch it are being removed. The industry is explicitly flattening middle management, with AI cited as having collapsed the overhead of being in the details. Layers are distortion machines — every interface between a decision-maker and the work softens the truth — so removing them removes distortion and the routing function at once. Nothing automatically replaces the routing, so the gap between reported status and user reality widens before it narrows unless someone instruments it deliberately.

The market started paying for precision instead of volume

Pointer's most useful observation is hiding in an advertisement: CodeRabbit bought the presenting-sponsor slot twice in one issue to run the line "Fewer Findings But All Of Them Real. Only Real Risks Should Reach Your Queue." A vendor spends acquisition money selling less output only when alert fatigue has become the dominant sales objection. The claims underneath are unbenchmarked; the positioning is the intelligence. If a feature you own surfaces findings, matches, suggestions or recommendations and its success metric counts what it surfaced, you are building to last quarter's buying criteria. Replace it with a false-positive ceiling, a suppression rate, and action-taken rate per surfaced item.

Computerworld supplies the agent version of the same defect, and it is the most useful reframe available for an exec review: McDonald's drive-thru AI failed because it did not know it was making mistakes, not because it made too many. A 95%-accurate agent that barrels confidently through the remaining 5% is worse in production than an 85%-accurate agent that raises its hand. That converts into three PRD gates most agent specs lack — calibrated confidence threshold, abstention and escalation rate target, and maximum time-to-human-handoff — plus one red-team output that matters more than accuracy: a count of silent errors, cases where the agent was wrong and proceeded anyway. The McDonald's material is a design principle to adopt, not a benchmark to cite; no error rates were published.

The depth number to anchor your targets on

When someone proposes an engagement target that assumes users will live inside your agent, The Information's figures on Meta's assistant are the cheapest available reality check: 500,000+ triers, 250,000+ daily actives and 2M+ prompts in roughly one week. That is a ~50% trier-to-daily-active ratio — a genuinely strong stickiness bar for a distribution-backed launch — against only ~1.1 prompts per active user per day. Even the best distribution in consumer software did not produce deep multi-turn daily habit in week one. These are unaudited internal figures one week post-launch, so they measure novelty-adjacent behavior, not durable retention. Use the first number as ambition and the second as the ceiling on any depth assumption in your model.

Why the justification arrives pre-built

One more trap worth screening for, from Buffett's 1989 letter as applied by an engineering leader who watched it at Shopify: any craving of the leader, however foolish, gets quickly supported by detailed rate-of-return and strategic studies prepared by the troops. The estimated damage from a single passing executive comment is six months of unnecessary work, and it defeats normal intake filters because it arrives as analysis, not as a whim. The only filter that catches it is provenance — who originated the idea versus who originated the justification.

What to do

  1. Audit findability on every feature shipped in the last two quarters this sprint: open each in the product's default state as a new user, classify findable / conditionally findable / hidden, and attach a fix ticket or an explicit kill decision to every hidden item.

  2. Replace volume-shaped success metrics on every AI output surface this sprint with a false-positive ceiling and an action-taken rate per surfaced item, and extend the definition of done to five rungs with an owner and telemetry source per rung.

  3. Add calibrated confidence, an abstention/escalation rate target and a maximum time-to-human-handoff as blocking criteria in every agent PRD before the next planning review, and run a red team whose only output is a silent-error count.

The bottom line

These items converge on one defect: the numbers your organization trusts — the vendor's rate card, the release dashboard, the leaderboard score — all describe supply, and none describe what a customer actually received. That breaks the two habits most AI roadmaps still run on: pricing features off published per-unit rates, and declaring them done at rollout. Both numbers drift in the direction that quietly costs you margin and adoption, and the drift stays invisible until finance or a renewal surfaces it. Pick your highest-traffic AI surface and make one owner publish two figures before the next planning review: what a single completed outcome costs, and what share of users reach it unaided.