The $2.6M Line Item That Repriced Your Model Budget
Two things landed in the same week: the top open-weights slot changed hands on post-training spend, and the largest enterprise software vendor made its own model layer rotatable in public.
What Xiaomi actually scaled
Pretraining compute did not move on this run. What moved sits downstream of it. Batch size and throughput: 1,568 samples per update, fully asynchronous, context up to one million tokens, 3.5–3.7 billion tokens per step. Task and environment diversity: coding, general agent, visual and cyber tasks mixed across harnesses, so gains in one capability reinforce the others. Grader compute: relative in-group comparison, which produced denser long-horizon reward and, notably, shorter solution paths per task.
The release boundary is the more instructive part. Roughly 7,000 environments, recipes and harnesses went out, covering coding, vulnerability reproduction, knowledge work, web development, even music scoring, while the complete task datasets stayed inside. Hugging Face has argued publicly that high-quality reinforcement-learning environments matter the way pretraining corpora did in the last cycle. This run is strong evidence for that view, and it points at an asset class most enterprises do not treat as one: internal workflows whose outcomes are programmatically verifiable, such as code review histories, support resolution traces, document decisions.
The demand side repriced as well
Microsoft stopped anchoring its enterprise AI stack on one lab. It now runs a single multi-model harness across Copilot, GitHub and security, with its own MAI models as default and GPT, Anthropic and customer-fine-tuned open weights all rotatable. Nadella told Stratechery the Anthropic anchor was "only a temporary state of affairs." Anthropic, in parallel, made a one-month data retention window the price of frontier access through Fable. Enterprises refused, adoption stayed low, and the requirement was removed in Fable 5.1.
A frontier model lost enterprise adoption over a retention window rather than a benchmark result. Capability is no longer the scarce input in the purchase decision.
The two arguments part company here
| Candidate for the new moat | Evidence behind it | What you must own to hold it |
|---|---|---|
| Reinforcement-learning environments and graders | Top open-weights position won on post-training; environments released, task sets withheld | Verifiable internal workflows, graded and version-controlled |
| Accumulated user context and memory | Meta's Muse judged the best personal agent while running on a model explicitly not state of the art | Workflow state and memory a customer cannot export |
| The harness itself | Claude Code and Codex show weak lock-in because artifacts live in GitHub and travel freely | Orchestration, tools and evals behind your own interface |
Both agree the weights are not the asset. They disagree on the replacement, and the disagreement has a budget line: the environments thesis funds a platform team, the context thesis funds product roadmap. Both are cheap to test this quarter and expensive to defer, because deferring leaves vendor choice standing in as the answer.
Where the evidence is thinner
- The headline figure is partial. It covers the final reinforcement-learning run, not pretraining a 1.02-trillion-parameter base. The defensible version is narrower and still valuable: with a strong base model in hand, frontier-adjacent capability is a single-digit-millions post-training problem.
- Top of open is not parity. A credible counter-read holds that leading closed and internal models operate a tier above anything on a public index.
- The scoreboards contradict each other. Grok 4.7 scored 56 on Artificial Analysis's Coding Agent Index while dropping five points to #24 on Vals. A frozen internal eval set is the only binding evidence in the room.
One procurement detail carries real exposure. Qwen-Image-2.1 drew an X post saying outputs were unrestricted while non-commercial language remained in the license text. The canonical license file governs in that conflict, and the licensee who relied on the post is the party carrying the risk once a customer builds on the output.
What to do
Commission a two-week production bake-off of the leading MIT-licensed open weights against your incumbent closed API on your three highest-token workloads, scored on cost per task, quality parity and latency.
Re-open model vendor terms this month on zero data retention, portability guarantees and benchmark-linked price step-downs, citing the Fable 5.1 retention reversal as precedent.
Make license provenance a hard procurement gate this quarter — canonical license file only — before any open-weight model enters a customer-facing path.