Open Weights Hit Frontier Parity — Reprice Before Your Margin Model Calcifies
The leaderboard flip is visible; the invisible part is that frontier-class quality's price collapsed while your unit economics sat untouched since 2024.
What actually broke
The headline is the leaderboard flip, but the number to internalize is the slope of the cost curve underneath it. GPT-4-class inference fell from $20 to $0.40 per million tokens in 36 months — a 50x drop — and open weights now route roughly a third of all OpenRouter traffic. Kimi K3 ranking first in six of seven frontend domains is the visible tip; the price of that quality collapsed while you weren't re-running the math.
Kimi K3 also ships two efficiency signals that change your evaluation method: it activates just 1.8% of experts per token and burns 21% fewer output tokens than its predecessor. Sticker price per token is now misleading — cost-per-completed-task is the real unit. A cheaper-looking model that loops or over-generates can cost more per finished job than a pricier one.
Where the sources diverge
Everyone agrees capability commoditized. They split on whether you can ship it. The optimistic read: weights publish under modified MIT on July 27 at $3/$15 per million tokens, self-hostable, no lock-in. The sober read is the use-vs-ship gap — 79% of developers use open models but only 51% ship them, widening to 57% versus 73% for closed at enterprise scale. That gap is operational, not quality: Kimi Delta Attention breaks prefix caching and needs 64+ accelerator supernodes, and Moonshot's 2.5x scaling-efficiency claim is self-reported and unverified.
Capability is now a purchased input. The question is no longer 'which model is best' but 'can my team operate the cheap one in production.'
The smart move is a bake-off, not a migration. Benchmark Kimi K3 against your current provider on your actual top workloads and measure cost-per-completed-task plus real deployment cost — then decide. Features you killed on unit economics in 2024 deserve a second pass against open-weight baselines; the margin math has moved by an order of magnitude.
What to do
Run a Kimi K3 bake-off this sprint against your current closed provider on your top-3 production workloads, measuring cost-per-completed-task — not per-token — once weights ship July 27.
Reprice every AI feature in the backlog against open-weight baselines by end of quarter and re-open any feature killed on cost in 2024.