The Open-Weight Discount Has a Reversibility Clause
The savings are real; the terms are not. Two sentences from Moonshot's founder turn every Chinese-weight dependency into a pricing and continuity exposure no renewal cycle has modeled.
What the vendor said out loud
Moonshot founder Yang Zhilin framed the open release in language worth keeping in a board pack. The company open-sourced because "right now, globally, we're not fully in the lead yet". That restates his earlier position that "the leader won't open-source; only the laggards will." Per ChinAI's translation of the Chinese-language coverage, he declined to commit to open source exclusively and reserved closed carve-outs for corporate deals. A supplier is describing its own openness as a catch-up tactic with an expiry date it alone controls.
The second disclosure is operational, not philosophical. Moonshot suspended subscriptions and imposed usage limits within 48 hours of launch because demand outran its compute. Winning a leaderboard and serving it are separate problems now. Availability risk on a single Chinese frontier vendor is concrete, not hypothetical.
Free weights are not free inference
Two of the accounts reviewed here look contradictory and are not. One prices the license at zero: 2.8 trillion parameters, a 1M-token context window, native vision, weights you can download. The other prices the hosted API at $2.30 per million blended tokens. Both are accurate. The distance between them is 1.4TB of MXFP4 weights and trillion-scale mixture-of-experts serving that somebody in the organization has to run, page and page again. Self-hosting converts a per-token bill into an on-call liability. That is the number that belongs in the comparison.
The tier structure is the part that matters
| Model | Blended $/M | Strategic read |
|---|---|---|
| Kimi K3 | $2.30 | Premium anchor; frontier parity converted to margin |
| Qwen3.7 Max | $1.40 | Best-placed premium-but-cheaper alternative |
| GLM-5.2 | $0.90 | Mid-tier, carries a Seoul-commitment disclosure gap |
| MiniMax M3 | $0.22 | Volume displacement at the bottom |
| DeepSeek V4 Pro | $0.18 | Default fallback for non-frontier work |
A 13x band separates the floor from the ceiling. The Chinese market is tiering, not commoditizing, and Moonshot's stated intent is to drag the sector "beyond cutthroat price competition toward value monetization." Any 2027 cost model that extrapolates a race to zero is extrapolating a strategy the leading vendor has publicly abandoned.
Three exposures sitting under the discount
- Pricing. The vendor has announced its direction of travel is up, and open weights give it no obligation to warn before the hosted tier follows.
- Policy. Commerce has opened an inquiry into Chinese developers distilling US model outputs with sanctions on the table, and Washington reportedly favors targeted restrictions on specific models. That is the worst outcome for anyone who hard-coded one. A White House official has publicly claimed Moonshot trained on Anthropic's intellectual property.
- Governance. Moonshot's own blog discloses K3's "excessive proactivity": unilateral decisions under ambiguous intent. It pushes mitigation onto the deployer through system prompts. Concordia separately finds zero safety papers from DeepSeek, Moonshot and MiniMax since January 2025.
The honest counterweight comes from Casey Newton's reporting. When Hugging Face was under live attack it could not use US frontier models to defend itself, because release conditions cap their cyber capability. It defended with Chinese open weights. The models under sanctions review are, in at least one documented case, the most effective defensive tools on the market.
Capturing the savings is fine. Making a model that Washington is actively considering restricting into a load-bearing dependency is not.
The disciplined posture is a cap, not a ban: a share of inference spend the business could absorb losing outright, a tested substitute behind an abstraction layer, and a switch cost already priced rather than one discovered under a deadline.
What to do
Set a named ceiling on Chinese open-weight dependency as a share of inference spend, signed off by the executive team this quarter
Commission a costed 90-day migration runbook with one pre-qualified substitute per critical workload, tested before the next renewal cycle
Package your own safety documentation, audit trails and bounded-agent behaviour as a priced differentiator in regulated-vertical deals this quarter