The Moat That Lasted a Day, and the One OpenAI Is Building Instead
Replication speed has reset what model access is worth, and the labs' answer is to contest a layer buyers cannot benchmark or price: regulatory access.
What made Astra credible was not the proofs
The ten results shipped with Lean 4 machine-checkable certificates, proof files a computer can verify without trusting the model that wrote them, at a total inference cost of roughly $2,000. Columbia's Henry Yuen vouched independently for the significance of the work. The reasoning replicated in a day. The certificate survived. Generation commoditizes; proof does not.
Pricing points the same way. OpenAI cut Luna's API price by 80% three weeks after launch and took Terra down 20%. Anthropic shipped Opus 5 at half the price of Fable 5. Artificial Analysis clocks DeepSeek's V4-Flash at roughly three cents per benchmark task against $1.86 for GPT-5.6 Sol and $3.15 for Claude Fable 5, level with Gemini 3.6 Flash on measured intelligence. Alibaba listed Qwen3.8-Max at $2/$6 per million tokens against Moonshot's $3/$15, downloadable weights a week behind. The cuts land before most enterprises have begun switching, which reads as preemptive defense rather than demand-driven price discovery.
Where the sources genuinely disagree
The cheap-inference consensus has one credible dissent, and it changes the arithmetic. V4-Flash's default reasoning setting produces code that is not shippable. Usable output requires reasoning set to high, which multiplies token consumption roughly six times. Delivered price is nearer $0.84-equivalent than $0.14, before the 167GB of memory needed to self-host it. Anyone re-baselining COGS or opening a renegotiation on the headline number is arguing from a figure off by a factor of six.
The demand side corroborates rather than contradicts. Token prices fell more than 95% in three years while enterprise LLM spend more than doubled in six months to $8.4 billion. Cheaper calls invited more calls. The only denominator that survives both facts is cost per accepted output at shipping quality.
The counter-move is regulatory, not technical
OpenAI previewed Astra to policymakers in Washington before it showed the system to customers, positioned as the first model submitted through the administration's pre-release framework, which is being finalized on a self-imposed deadline. Whoever goes first authors the documentation expectations and the evaluation batteries everyone smaller inherits.
The mirror image landed in Europe on 2 August, when the EU AI Office gained the power to inspect models before launch and block them from the market, with fines up to €15 million or 3% of global turnover and obstruction independently fineable. The Commission opened talks with OpenAI and Anthropic within days. The vendor's regulatory calendar now sits inside the European launch date, which makes single-vendor model dependency a market-access risk almost nobody has priced.
The frontier is still worth paying for where quality genuinely dominates. Everything beneath it now has four credible suppliers and no lock-in.
One escape hatch is already closed, and it closed on evidence. Routing is not the answer: Manifest ran an LLM router in production across 7,000 users for four months and chose to deprecate it. Abstraction is a commercial necessity, not a business.
What to do
Run a cost-per-accepted-output bake-off within 30 days across your three highest-volume agentic workloads, pricing DeepSeek V4-Flash at high reasoning against your incumbent before anyone claims a switching saving
Commission an in-scope determination on the federal pre-release submission framework and name your EU AI Act representative before your next European launch date is committed
Cap any single model provider at 60% of inference spend this quarter and prove a primary-model swap in staging inside two weeks with no application change