Product & Strategy

The Product Desk

The Signal

Canva cut its growth outlook by a third because its AI features worked too well.

Demand "significantly exceeded expectations," and serving it pulled guidance down to 20% growth; Figma's cash margin fell from 27% to 14% in a single quarter. That pairing has been the quiet story all year — inference cost lands on the income statement long before it lands on a roadmap. The counterweight is the number worth stealing: cost per task down nearly 90% since April, earned by re-architecting delivery rather than shipping less. The forcing function before any AI feature of yours reaches GA is whether you can state its cost per task today. If you can't, usage growth is a liability line, not a win.

In Play

  1. AI Serving Cost Reached Revenue Guidance

    Canva cut its expected revenue growth by a third, to 20%, because delivering its AI features cost more than planned. Figma's free-cash-flow margin fell from 27% in Q1 to 14% in Q2, and PitchBook analyst Derek Hernandez calls the pair the clearest sign inference cost is eroding software economics. Inference is now a guidance-level variable rather than a COGS footnote, which puts a cost-per-task number on your launch gate.

    Ask Clarity
    Try
  2. Cheapest Token Stopped Meaning Cheapest Task

    DeepSeek shipped V4-Pro on Wednesday at $0.435 per million input tokens and $0.87 per million output, against Claude Opus 5's $5 and $25. The Information reported the same week that premium OpenAI and Anthropic models can win on effective cost — the performance-adjusted price of finishing a job. Both findings can hold, and only one denominator survives contact with your P&L: cost per successfully completed task. DeepSeek has also pre-announced a significant API price increase.

    Ask Clarity
    Try
  3. Governance And Context Gate AI Ship Dates

    Cloudera reported on Aug 11 that 95% of enterprises delayed or canceled AI projects in the past year, with governance and compliance among the biggest blockers and about 75% saying AI made data governance more complex. A separate survey of 101 enterprises found bad context, not weak models, behind confidently wrong agent answers; teams with a governed semantic layer caught roughly twice as many of those failures. Your ship date is gated by context and permissions, so fund the context layer before the model upgrade.

    Ask Clarity
    Try
  4. AI Features Got A Public Price Tag

    Google raised every Pixel 11 tier by $100 — $899 base, $1,099 Pro, $1,899 foldable — attributing it to memory chip prices while shipping minimal hardware change. Morning Brew reads the delta as what buyers are being asked to pay for Gemini Intelligence software, and Apple has already repriced MacBook and iPad on the same input costs, per Tim Cook. That hands your AI pricing debate a public willingness-to-pay anchor. Wendy's six straight quarters of same-store sales declines says consumers are still trading down.

    Ask Clarity
    Try
  5. A Court May Write Your Age-Assurance Spec

    A coalition of state attorneys general takes Meta to trial in Oakland later this month over practices affecting young users, seeking court-ordered changes to age verification and data collection rather than damages alone, per Bloomberg's Riley Griffin. Remedies imposed on the largest player tend to become the practical compliance baseline for everyone shipping to minors, well before any statute lands. If your product has under-18 users, age assurance becomes a platform-identity project rather than a consent checkbox.

    Ask Clarity
    Try

Deep Dives

The Guidance Cut That Came From Winning

Two adoption forecasts point in opposite directions, and the pair decides which number belongs on your GA gate rather than in your post-mortem.

The number under the guidance cut

Someone opened Canva in April, tried the new AI feature, kept the output, and used it on the next document too. No complaint, no support ticket. That is the entire user story sitting underneath a guidance cut. Canva's disclosure carries the recovery figure that matters more to a roadmap than the miss does: cost per task is down nearly 90% since Canva AI 2.0 launched in April, achieved by re-architecting how the features are delivered rather than by shipping fewer of them. That is a public benchmark to carry into an engineering review. It says first-generation AI features were badly unoptimized as a class, and that the headroom between a naive implementation and a tuned one runs close to an order of magnitude.

The sequence is the component worth copying. Canva slowed the rollout first, rebuilt the architecture underneath, and absorbed the growth cut in public while the work happened. Melanie Perkins' explanation was that user demand for the AI features "significantly exceeded expectations." Nothing failed a quality bar. Nobody was indifferent. Adoption beat plan, and beating plan is what broke the model.


Two adoption forecasts, missed in opposite directions

Set that against The Information's reporting that ChatGPT is nearing 1 billion weekly active users, seven months after OpenAI's own target. The most distributed consumer AI product in history missed its own adoption forecast by more than two quarters. Canva's embedded AI features overshot theirs badly enough to move guidance.

Both numbers hold, and they are not in tension. They are two error modes with two different owners. Standalone AI products overestimate the pull of a new surface users have to travel to. AI features embedded in a workflow users already live in underestimate the volume those users will push through it. Only the second error lands in COGS. It lands there while the launch is still being celebrated.

SignalNumberWhat it re-prices
Canva revenue growth outlookCut by a third, to 20%Your AI business case at optimistic adoption
Figma free cash flow margin27% (Q1) to 14% (Q2)Inference as a margin line, not a footnote
Canva cost per task since AprilDown ~90%Optimization you can fund instead of descoping
ChatGPT weekly actives~1B, 7 months lateThe adoption curve in every AI business case

Most teams cannot produce the number

Here is the gap that makes this urgent rather than merely interesting. VentureBeat's reporting puts two-thirds of enterprises running AI in production, many with no visibility into what that infrastructure costs or how well it is utilized. They bought AI compute for speed and are flying blind on price. Canva has now demonstrated in public that this number decides guidance. Most product organizations cannot generate it for one feature, let alone at forecast adoption.

One caution before the benchmark gets pasted into a deck: a ~90% reduction from a first-generation implementation says more about the starting point than the ceiling. Treat it as evidence that a teardown is fundable, not as a target owed to a CFO.

For an AI feature, the optimistic adoption scenario and the worst-case margin scenario are the same scenario.

The move

The forecast field that catches this is not the base case. Model the P95 adoption curve, multiply it by measured cost per task, and put that product on the GA gate. Then run the descoping list backwards, because features parked on gross-margin grounds were priced against an unoptimized implementation, and optimization is now demonstrably cheaper than the feature cut. Two axes for the review: cost per task measured or assumed, adoption modeled at base case or at P95. One cell survives a launch that works.

What to do

  1. Add a mandatory 'cost per task at P95 adoption' field to your AI feature PRD template and make it a GA gate before your next launch review.

  2. Commission an inference teardown of your top three shipped AI features this sprint with an explicit 50% cost-per-task reduction target.

  3. Re-baseline your AI adoption forecast against the seven-month miss and present the revised curve at the next planning cycle rather than after you miss.

DeepSeek Reset The Floor And Pre-Announced A Price Hike

Cheapest per token and cheapest per finished job are now different vendors, which turns what looks like a procurement call into a design decision your architecture may not support.

The flagship lost to its own smaller model

An engineer runs the same coding task through the new flagship and through the cheaper model already in production, and the cheaper one wins. That is the launch detail worth acting on. Testers reported V4-Pro losing continuity in its reasoning, performing weakly on image tasks, and on certain coding tasks doing worse than DeepSeek's own smaller V4-Flash, the late-July release that got popular precisely for strong performance at low operating cost. On the Vals AI open-source leaderboard V4-Pro sits second, behind Moonshot's Kimi K3 at $3 input and $15 output per million tokens. Bigger stopped being better inside one lab's own lineup.

Provenance caution: those quality reports come from engineer testers on social media, not controlled evaluation. Run your own evals on your own tasks before a roadmap item depends on the cheap tier.


Where the two price signals actually conflict

The Information reported on Aug 14 an industry analysis finding that premium Anthropic and OpenAI models can beat Chinese rivals on effective cost, the performance-adjusted total cost of finishing a job, including retries and reasoning overhead. That sits awkwardly beside a 1/29th output price. The conflict dissolves only when the denominator changes.

Separate the thing being priced from the thing being bought. A model that costs 29 times less per token and fails three times as often is not cheaper. A model that costs more per token and succeeds on the first attempt can win on cost per resolved ticket while losing every price-per-token comparison run to date. The two reports are not disagreeing about vendors. They are disagreeing about the unit. The forcing function is one number, measured on the workload already in production: cost per resolved ticket, not cost per million tokens.

Weigh the study accordingly: its sponsor and methodology were not disclosed, and the finding favors the two vendors most exposed to Chinese price undercutting. Useful as a measurement method, weak as evidence in a PRD.

ModelInput / 1MOutput / 1MPositionHow to use it
DeepSeek V4-Pro$0.435$0.87#2 open source (Vals AI)High-volume work where a wrong answer is recoverable
DeepSeek V4-FlashBelow V4-ProBelow V4-ProLate-July cost/perf hitDefault candidate; reportedly beats V4-Pro on some coding
Moonshot Kimi K3$3$15#1 open source (Vals AI)Quality leader in open weights, 7-17x the floor
Claude Opus 5$5$25Premium frontierEscalation tier where an error is expensive

Two reasons today's floor is unsafe to price against

First, DeepSeek has already said it plans a significant increase in overall API pricing while pricing V4-Pro at loss-leader levels. Land grab now, monetization step scheduled later. A pricing page, a customer commitment, or a business case built on $0.435 underwrites somebody else's share-buying strategy with a gross margin that is not theirs.

Second, leverage is arriving from the other direction. Meta has put thousands of engineers on closing the coding-model gap with the stated goal of reducing reliance on OpenAI and Anthropic. Nvidia is reportedly building best-in-world open-source models, and Applied Compute is in talks to roughly double its valuation to about $3 billion from $1.5 billion specifically on open-source demand, which is a reported financing discussion and not a closed price. Two years of duopoly pricing is becoming a three-way race with a credible self-host path behind it. The buyer with unlimited engineers refuses single-vendor dependency, and contract renewals inherit that argument.

The architecture that wins routes each task to the cheapest model that still succeeds, and you cannot make that call without measuring success.

What to do

  1. Rewrite your model-selection scorecard this sprint so the primary metric is cost per successfully completed task, including retries, tool-call loops and human escalation, with $/1M tokens demoted to a supporting column.

  2. Write a one-page routing-layer PRD by month end: cheap model as default, escalation on validation failure, per-workload provider config, and a hard per-request cost ceiling.

  3. Re-open your model vendor contract this quarter with two asks: a 12-24 month price ceiling and committed throughput for launch windows.

The Agent Control Plane Enterprise Buyers Will Ask For Next

Four vendors in unrelated categories converged on the same requirement, and a documented case of agent harm shows what shipping without it costs.

An IT lead goes looking for a list of the agents running inside her environment. She finds service accounts, API keys, and a directory of humans. None of those is the list she wants. She is not disorganized. The inventory she needs was never a field in any system she owns. That gap is what JumpCloud's Q3 2026 research, based on 800 IT leaders, points at: agents are already running critical workflows while most IT teams have no systematic way to see, own, or shut them down. Separate the thing being pitched from the thing being done. The pitch is agent governance. The thing being done is an admin trying to answer two questions during an audit: what is running, and who is accountable when it misbehaves. Those are not the same product. The first is an inventory problem. The second is an org chart problem that no dashboard resolves. Positioning has converged with unusual precision across categories that share no market. That convergence is the most interesting line in the research and the least reliable signal in it. Vendors converge on language well before customers converge on behavior. Everyone writing the same sentence about agent sprawl tells you what the roadmap slides will say next quarter. It does not tell you that an agent inventory exists in production anywhere. The number that would settle this is not how many leaders report agents in critical workflows. It is how many can produce the list when asked, and what happens to the workflow when one entry on that list gets revoked. The research says the systematic capability is absent for most of those leaders. Treat that as the baseline condition, not as news. The diagnostic worth running is a grid. One axis: can the team enumerate the agents acting in production. The other axis: can it revoke one without asking the team that built it for permission. Enumerate and revoke is governance. Enumerate without revoke is a report with good formatting. Revoke without enumerate is luck, and it holds until the first agent nobody remembered breaks a workflow that finance depends on. The research places most teams in the cell with neither capability. The forcing function for this week is small and unpleasant. Pick one critical workflow an agent already touches. Name its owner in writing. If the answer comes back as a team rather than a person, the visibility problem is an ownership problem wearing a dashboard, and buying a platform first will produce a longer list with the same unanswered question at the end of it. Build the ownership field before shopping for the tool that promises to populate it.

What to do

  1. Open an enterprise-readiness epic this quarter covering agent inventory, per-agent identity, one-click kill-switch and per-agent spend attribution.

  2. Classify every agent action in your product as reversible, irreversible or third-party-affecting by quarter end, and require explicit confirmation plus an audit record for the last two.

  3. Add a harness section to your trust center and security questionnaire responses this sprint: every tool the agent can invoke, argument validation at the tool-call boundary, and per-call logging with caller identity.

The bottom line

Today's items rhyme in one uncomfortable way: the hard part of an AI product has moved from getting the capability to affording it, proving it, and controlling it once customers actually use it. That kills two comfortable assumptions — that a feature which survives a demo is a feature you can ship, and that a cheaper provider is a margin strategy. Capability is the commodity now; the constraint is what you can measure and what you can switch off. Pick the AI surface with your fastest-growing usage and write down two numbers this week: what one successful outcome costs, and how fast someone can stop it.