AT&T Priced The Quality You Give Up For Cheaper Models
Routing stopped being a research spike: operator numbers, a flat-rate counterattack from Replit, and three published efficiency levers together reset what your margin defense has to contain.
What the 56% is actually measuring
A developer inside Ask AT&T opens the model picker, sees GitHub Copilot, Devin, Claude Code and Codex, and takes whichever one she used last week. She is not optimizing anything. Her spending is capped, so the choice that feels like a capability decision was already made as a budget decision. AT&T tiers by task complexity, not by user: frontier models generate code, cheap open models summarize code that was already written. Premium AI tooling ended up as a menu of interchangeable options competing under a ceiling someone is actively lowering. That lands on the pricing page, not the architecture diagram.
Mark Austin, the VP running AI for AT&T's 100,000 employees, told The Information the open-to-frontier capability gap runs six to 10 months and appears to be narrowing, and that AT&T is repatriating inference onto its own Nvidia and AMD hardware because it beats renting cloud capacity. That is a roadmap clock, not a benchmark note. A capability that needs a frontier model today plausibly runs economically on open weights in two to three quarters. Model-brand positioning is depreciating collateral.
The cost-adjusted frontier moved underneath the labs
Gemini 3.7 Flash posted 84.6% on ARC-AGI-2 at $0.25 per task and 95.5% on ARC-AGI-1 at $0.12. GLM-5.3 Max reached 1597 points in Code Arena WebDev at $3.65 per million tokens. The caveat travels with the numbers. Practitioners flagged the leading community coding eval as saturated, with all models bunched at the top, and Qwen3.8-27B regressed against its predecessor on offline factual recall while improving at tool use. Separate the thing being pitched from the thing being measured. A public leaderboard can no longer carry a model decision. A private golden set can.
The meter, not the price, is the objection
Replit folded usage costs into its flat monthly subscription, claiming subscribers can create up to 30 times more than before on OpenAI's GPT-5.6 Luna. At the other end of the same market, a $200-a-month plan is reportedly exhaustible in one heavy Codex day, with consumption continuing past the stated cap. Sources disagree here, and the disagreement is useful. One camp reads metered AI pricing as a competitive liability to remove. The other says reprice off token consumption rather than seats and define an overage mechanic. Both demands resolve to the same missing artifact: the P95 heavy user's true inference cost measured against a frozen quality baseline. Replit's headline capacity claim is also bound to one supplier's price sheet, which is gross margin outsourced.
Three levers that cost days, not quarters
- Semantic caching benchmarked at a 57.1% hit rate, 55.7% fewer tokens and roughly 15 ms hit latency.
- Gisting at about 40% lower end-to-end latency and 15% higher throughput, per Shopify's writeup.
- Markdown tool output instead of JSON, at roughly half the tokens for the same payload.
One trap sits underneath all of it. Vendor prompt caching matches on an exact prefix, byte for byte: reorder two cached policy documents and both become misses, and the observed production pattern is a small number of blocks serving nearly all hits. A savings claim sourced from an aggregate cache-hit dashboard is a projection, not a result. The forcing function for this sprint is narrow. Take the two highest-volume AI calls, measure cost per successful task against the frozen golden set, and route to open weights anything that clears it. What fails becomes a frontier line item with a name attached.
Every AI feature you shipped without a routing layer is now provably about twice as expensive as it needs to be, and your enterprise buyer has the receipt.
What to do
Instrument cost-per-successful-task plus a frozen quality baseline for every AI feature, broken out by feature, task and model, before your next pricing review
Run a routing spike on your single lowest-complexity AI task within two weeks using a gateway plus an open-weight model, targeting a 40% cost cut at under 5% quality regression
Model your top-decile user's real inference cost against your subscription price this month and publish the fair-use ceiling where a flat tier stays margin-positive