Anthropic's Metering Shock: Your Claude Cost Model Is Wrong in 30 Days
What Changed
Anthropic converted Claude subscriptions from flat-rate developer access into dollar-matched API credits. Every Agent SDK call, GitHub Action, claude-p pipeline, and third-party harness job (Conductor, Zed, OpenCode, T3 Code) now meters tokens at list price. The 70-90% effective discount alternative-harness users had been arbitraging is gone. Starting June 15, third-party tool usage gets a separate credit bucket. No subsidized tokens, no rollover, overflow at API rates.
Same day, Sam Altman posted a 2-month-free Codex enterprise switch promo. That timing is not coincidence. Ramp's April data shows Anthropic edging OpenAI 34.4% vs 32.3% in business spend, the first lead change. Anthropic has hired a CFO and is targeting an October IPO. Margin-per-token is now a board metric.
The Capacity Context
Dario Amodei conceded they planned for 10x growth and got 80x in revenue and usage. The fix is leasing xAI's entire Colossus 1 cluster — 220,000+ GPUs (H100, H200, GB200). Rate limits are loosening: Claude Code 5-hour limits doubling, peak-hours throttle removed, Opus API limits "substantially raised."
Any Claude benchmark from before May 7 is stale. The serving conditions your eval harness measured are about to shift again.
ServiceNow's CDIO has already burned through the full-year Claude budget by May. National Life Group's CIO calls Claude "great for consumer usage but not great for companies" wanting per-user monitoring. The thing the leaderboard score doesn't tell you: Anthropic provides no native per-user telemetry and no SLAs on latency or availability.
What Sources Agree On
Nine independent sources converged on the same conclusion: single-provider Claude dependency is now the highest-risk default on the table. The disagreement is on severity. Some frame it as a transient capacity issue that Colossus resolves in weeks. Others frame it as structural IPO-driven margin extraction. Both can be true simultaneously, and the cautious read is to plan for both.
The Contradiction Worth Noting
Anthropic is simultaneously the fastest-growing enterprise AI provider (ARR reportedly tripled from $9B to $30B+ in four months) and the one with the worst observability story for enterprise customers. Growth plus no telemetry plus no SLAs plus price increases produces the ServiceNow outcome at scale. The capacity numbers are correlated with the margin push. Causation runs through the IPO calendar.
What to do
Audit every Claude-backed workload (Agent SDK, GitHub Actions, batch evals, third-party IDEs) and project token burn under metered pricing by end of this sprint
Deploy an LLM gateway (LiteLLM/Portkey) with per-user, per-feature token tagging and daily budget alerts within 2 weeks
Run OpenAI's 2-month Codex enterprise switch promo as a controlled A/B on your top 3 Claude workloads
Avoid locking annual Anthropic contracts until post-Colossus serving stability is observable (target 6-8 weeks)