Product & Strategy

The Product Desk

The Signal

Rewriting SWE-bench repos without changing the code cut model success up to 14.4 points.

The smallest drop was 6 points, measured across four models, with input tokens up more than 2.5x on the same tasks. No model has trained on your private repo. The leaderboard therefore flatters both the success rate you budget for and the cost per task you end up paying.

In Play

  1. Opus 5.5: Cheaper Per Token, Unproven Per Task

    Across today's stories, producing work got cheaper while checking it did not. TheSequence and Simplifying AI report that Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens. Anthropic says typical workloads cost about 40% less than on Opus 5 and output runs more than 30% faster. Both outlets stress that these are vendor figures that nobody has independently verified yet. The per-token price is settled, but the saving per successful task on your own code is not. Re-score the AI features you shelved over margin or latency against your own eval, not the launch post.

    Ask Clarity
    Try
  2. SWE-bench Overstates Your Repo

    TheSequence covered the SchrodingerRepo paper, which found signs of leakage in more than 65% of SWE-bench Verified instances. Leakage means the test answers likely appeared in the models' training data. Researchers then rewrote repos without changing what the code does. First-try success (Pass@1) fell 6.0 to 14.4 points across four models, and input tokens rose more than 2.5x. No model has seen your private codebase, so vendor scores will overstate your agents' success rate and cost efficiency.

    Ask Clarity
    Try
  3. Agents Outrun Their Reviewers

    Three practitioners writing in Architecture Notes independently reached the same conclusion: multiple agents now change a system faster than anyone can review every diff. TheSequence reports that on Taste-Bench, a set of 502 questions posed at points where an agent must choose a path, the best model scored 59.7%. A bigger reasoning budget did not help. Your agent features need human checkpoints where the agent makes those choices, not just a review after the run ends.

    Ask Clarity
  4. Free Submissions Break Paid Intake

    Risky Business reports that Intel quietly switched its bug bounty to 'No bounty.' The program had paid up to $100,000 per confirmed report, and the change comes as AI-generated reports flood bounty programs across the industry. Separately, in February, fraudsters used an AI-cloned lawyer's voice to get €95M ($108M) out of Fideuram. After recoveries, the bank's net loss is €36M. Every funnel you run that pays for or trusts each submission, from support queues to voice approvals, was priced for human effort.

    Ask Clarity
    Try
  5. Gen Z Bets Its Investing Money

    Morning Brew cites a Betterment survey of 1,000 investors in which 52% of Gen Zers said they redirected money meant for investing into sports betting. Bank of America counted 22% more first-time bettors last football season than the season before. If you build money products for young users, your real rival for their spare cash may be FanDuel.

    Ask Clarity
    Try

Deep Dives

Opus 5.5's Discount Is Proven Per Token, Unproven Per Task

What you can actually save depends on success rates that no public leaderboard can measure for your code. The eval you build this sprint decides which shelved features come back.

Only one cost term is settled

Your re-score needs one number: cost per successful task. That is tokens per attempt, times price, divided by success rate. Anthropic's release fixes the price term and shifts the token term. Simplifying AI worked the arithmetic. The per-token price fell 20%, so a net saving near 40% implies about 25% fewer tokens per task (0.8 × 0.75 ≈ 0.6). That is a derived figure, not an Anthropic disclosure. It still matters for your dashboards. A dashboard that compares models on dollars per million tokens will miss roughly half of this gain, and every future efficiency gain like it.

No vendor can give you the success-rate term. The SchrodingerRepo researchers found that on rewritten repos, 81.6–83.6% of the extra actions agents took were exploration, meaning the model searching through code it had never seen. That searching is what pushed input tokens up. Your private codebase triggers the same behavior, because no model has trained on it.


What unfamiliar code does to a business case

TheSequence ran an illustration with the paper's GPT-5.4-mini result. On a rewritten repo, input tokens rose 2.5x and first-try success fell from 46.8% to 35.6%. Together, that means roughly 3.3x the input-token cost per success. Even after a 40% price cut, you are near 2x what a leaderboard-based projection implies. The math mixes models and setups, so read it as a direction, not a forecast.

This risk doesn't hit all your planning documents equally. A cost baseline pulled from production already reflects your own code. A business case built from vendor demos or SWE-bench scores does not, so that is the one most likely to overstate ROI.

Both outlets say the headline figures are self-reported, but they differ on how fast to move. Simplifying AI treats Opus 5.5 as a model swap an existing Anthropic customer can make this week. It also suggests a routing pattern: default to Opus 5.5 and send only the hardest requests to Fable 5.1. That works because Anthropic claims Fable 5.1-level quality on “most tasks.” Simplifying AI adds that neither “most tasks” nor “typical workloads” is defined. TheSequence puts more weight on checking results against your own code before you commit.


A cost lever you control

Salesforce AI Research built JIT Mem, which stores raw records of past agent runs and assembles task-specific context only when a task needs it, instead of summarizing memory in advance. In the ALFWorld, WebShop and τ²-bench test environments, it cut input tokens 50.3–56.3% and steps 28.4–31.4%. It also beat memory systems that summarize in advance by up to 16.3 success points. Even an untrained Gemini curator scored 61.0, against 41.0 for SkillOS. These are benchmark environments, not production codebases. Still, JIT Mem targets the same exploration overhead the rewrite study exposed, and memory design is a decision your team owns.

The smart move

Build the eval before the business case. Draw 50–100 tasks from your own repos and workflows. Add variants that change how the code looks but not what it does, such as renamed namespaces and reordered files. They show how much a score depends on code the model has seen before. Then re-score the shelved backlog from those results, not from the launch post. Finally, clean up your own claims. Technical buyers can now easily challenge unqualified SWE-bench Verified scores in your PRDs or sales decks.

What to do

  1. Build a 50–100 task private eval from your own repos and workflows this sprint. Include variants that preserve what the code does, and compare Opus 5.5 with your current model on success rate and tokens per successful task.

  2. Re-score every AI feature you shelved for cost or latency using cost per successful task from that eval, and bring the top three into Q4 planning before the roadmap locks.

  3. Audit PRDs, sales enablement and marketing pages this sprint for unqualified SWE-bench Verified claims. Add caveats or replace them with private-eval results.

AI-Drafted Specs and End-of-Run Reviews Both Miss the Fork

Agents make their costliest mistakes mid-run, at branch points. The artifacts PMs rely on cover the start and end of a run and leave the middle unwatched.

The spec finding aimed at your desk

Ayman Nadeem built Nuanced around persistent plans. He then found that long AI-generated specs added friction without improving understanding. His essay is titled “Plan Mode Is Dead,” but it argues against separating planning from building, not against planning itself. If your PRD process starts with “have the model draft the full spec,” this is a warning from someone who organized a whole product around plans: that step may be theater.

Fatih Arslan shows what survives. He keeps plans as living Markdown files in a shared Git repo, sorted into drafts, queued, active and completed folders. Long-running coordinator agents keep the index current. In effect, it is a Kanban board that humans and agents both read. Architecture Notes reconciles the two views: plans survive as a shared record of decisions and die as an upfront approval gate.


Why reviewing the finished run is too late

TheSequence's research roundup shows where oversight should move. Taste-Bench, from Microsoft and CityU Hong Kong, tests agents at decision forks. These are the moments where an agent picks an approach, expands scope or takes a destructive step. The longer the task, the harder the forks got. A related result is more encouraging. A distilled “advisor” model, meaning a smaller model trained to make judgment calls, advised an unchanged executor agent. With that advice, the executor's success rose from 14.6% to 33.7% on 41 held-out SWE-bench Pro tasks. Forty-one tasks is a small sample. The design lesson is that judgment at the fork is a separate component, and either a human or a small model can supply it.

WhatWorkedBench, from CMU and Tsinghua, adds a warning for products where agents analyze experiments. It covers 36 tasks and 1,248 configurations, and 35 involve sign-reversing interactions, where one setting flips the direction of another's effect. Agents found settings that worked but missed why they worked. The researchers paired agents with a Gaussian process, a statistical model that estimates effects from data. That raised effect recovery from 0.632 to 0.698 in one group of models and from 0.303 to 0.455 in another. If your product lets an agent explain A/B results or configuration choices, don't ship its “why” unaided.


What the oversight surface should contain

Brad Murry catalogued 13 orchestration patterns, ranging from pipelines and fan-out to supervisors and review loops. That gives your specs a shared vocabulary. Each pattern comes with use cases and failure modes. His design runs all of them on typed shared state with explicit transitions, which makes runs easier to inspect and recover. OpenMuse, an open-source agent covered by Simplifying AI, ships the checkpoint as a default: a separate gatekeeper asks for confirmation before any irreversible action. Simplifying AI notes that the gatekeeper protects users only if it correctly identifies which actions are irreversible.

Together, these point to three requirements most agent PRDs lack:

  1. A named pattern with its failure modes, so reviewers know what breakage looks like.
  2. A retry budget: the maximum attempts and spend before a human is pulled in. That is a product decision about cost and trust, not an implementation detail.
  3. A decision record for each run: what the agent decided, which alternatives it rejected and which systems it touched, readable without opening a single diff.

What to do

  1. Add a decision-record requirement to your next agent-feature PRD this sprint. Measure how long reviewers take to approve a run and what share of runs produce a complete decision summary.

  2. Before each agent workflow's next beta, specify human checkpoints at three points (approach choice, scope expansion and destructive actions) plus a retry budget in attempts and spend.

  3. Start a two-sprint experiment that replaces long AI-drafted specs with a one-page living plan kept in the repo. Track time from spec to first build and the number of clarification questions from engineering against your current baseline.

Intel's Dropped Bounty Previews What Free AI Output Does to Your Queues

Bug programs, voice approvals and other systems priced for human effort now absorb machine-made volume. The durable fix is requiring proof, not paying smaller rewards.

What is known, and what is inference

A side-channel researcher opened Intel's bounty page and found no payout figures. The last archived version showing them is dated September 13. Intel has refused to explain the change to Risky Business or to the hardware news sites that spotted it. The program ran on Intigriti for more than a decade and added large bounties after Spectre and Meltdown. The link to AI-generated report volume is Catalin Cimpanu's inference, not Intel's statement. The pressure itself is documented. Over the past year several tech giants have said AI-found bug reports clog their programs and drain triage budgets. AI pen-testing products such as PortSwigger's Burp AT will push that volume higher.

Separate the pitch from the act. The pitch is noise reduction; the act is removing the cash. Academics working on side-channel and transient-execution attacks had reported payouts in the tens of thousands of dollars. Those go with the noise. Cimpanu predicts other companies will cite AI volume to justify the free-report model of two decades ago. Treat that as informed opinion, not a trend line.


Every queue runs on the same economics

Swap “bug report” for any unit a product handles one at a time: support tickets, feedback forms, app reviews, referral sign-ups, job applications. Each assumed real human effort per submission. When a submission costs almost nothing to produce, the cost moves to whoever reviews it. Cheaper models keep pushing that production cost down.

Identity is the same problem with bigger losses

The Fideuram fraud shows the bill. In February, fraudsters impersonated Intesa Sanpaolo CEO Carlo Messina in messages to Paolo Molesini, then chairman of Fideuram, the bank's private banking arm. A cloned lawyer's voice then got an executive to approve transfers to accounts in China and Hong Kong. More than half the money has been recovered. €36M is still missing. Analyst Hayden McKenzie separately tracked a North Korean IT worker cell operating behind a US front company called MageHire; at least three women appeared on camera in interviews and meetings while male developers did the work off-screen.

Simplifying AI flagged VoiceStudio, a free tool that runs entirely on local hardware and handles zero-shot voice cloning, dubbing and transcription in 646 languages. AudioSeal watermarking is on by default, and anyone with the code can switch it off. Cloned voices will circulate unwatermarked.

A submission, a voice or a face now proves only that someone had access to cheap generation tools. Trust has to rest on something costing the sender real effort or access: a proof others can reproduce, a reputation history, or a cryptographic confirmation over a separate channel.

The smart move

Measure before changing incentives. The useful number is the AI-generated share of each queue, per queue. Then filter without taxing honest senders: reproducible proof of concept, reputation weighting, duplicate clustering, AI pre-triage. The forcing question is whether a change raises the cost of submitting or the cost of being trusted. Intel raised neither. It cut the reward, which cuts the good reports too and may push the best researchers toward exploit brokers.

What to do

  1. This sprint, measure the AI-generated share of submissions in every intake queue you own (support, feedback, reviews, referrals, bounty). Add proof requirements, duplicate clustering or AI pre-triage before changing any reward.

  2. This quarter, stop accepting voice or video as the only approval for above-threshold payments, admin role grants, account recovery and remote hiring. Require confirmation over a separate, cryptographically secured channel instead.

The bottom line

Today's stories share one asymmetry: producing work got cheaper, while confirming that the work is correct, safe or authentic did not. That breaks the habit of treating a model price cut as a feature price cut. The next budget line to grow is checking: evals, reviewers, triage queues and identity proof. Fund that layer before you reopen the backlog, and build one tool this week that measures success on your own work, so its results, not launch announcements, decide which shelved features return.