The Cheapest Frontier Model Is Also the One That Guesses Most
Two measurement houses graded the same release and disagreed on the headline, and which one you believe decides whether accuracy-critical features ship this quarter.
The mechanism behind the 50%
Start with what the model does, not what the benchmark says. Opus 5's hallucination rate did not climb because the model got worse. Artificial Analysis attributes the 14-point jump to a behavioral change: the model answers more often when it is unsure rather than declining. That turns a research metric into a design variable you own. If your surface needs abstention — "I don't know" as an acceptable output — you now have to build it yourself in prompting, retrieval grounding, or a post-check. The model has been tuned to volunteer.
Two scorecards, two headlines
Artificial Analysis and Epoch measured the same release and told different stories. Artificial Analysis has Opus 5 leading its Intelligence Index at 61 against Fable 5's 60 and GPT-5.6 Sol's 59, sharing first place for coding at 89% on Terminal-Bench v2.1, and posting 1720 Elo on AA-Briefcase — its head-to-head ranking for simulated office and agentic work — a 146-point lead over Fable 5. Epoch's numbers are flatter: software-engineering parity (SWE-ECI 161 versus 161) and a slight general-capability trail (ECI 159 versus 161).
Reconciled, the picture is precise rather than contradictory. Opus 5 buys genuine separation on agentic and coding work, near-parity on general capability, and a regression on factual answering. Fable 5 remains the most accurate of the three.
| Model | Capability | Coding / agentic | Reliability | Cost posture |
|---|---|---|---|---|
| Claude Opus 5 | 61 index / ECI 159 | Shared #1; 89% Terminal-Bench; 1720 Elo | 50% hallucination rate | ~½ Fable's price; -20% per task |
| Fable 5 | 60 index / ECI 161 | 1574 Elo AA-Briefcase | Best factual accuracy of the three | Baseline premium |
| GPT-5.6 Sol | 59 index | — | — | Efficiency parity with Opus 5 |
| GLM 5.2 (CompactifAI) | Sonnet 5 tier (claimed) | Drop-in for Cursor, n8n, LiteLLM | Parity claimed, unverified | $3.50 per 1M output tokens |
The bottom of the market filled in the same week
GLM 5.2, distributed through CompactifAI, claims Sonnet 5-tier quality at 65% below Sonnet 5's output-token price with no migration work for teams already on Cursor, n8n or LiteLLM. Put that next to Opus 5's repricing and the expensive habit becomes obvious: defaulting a single model across every surface. Output-heavy, low-stakes workloads are overpaying at premium tiers. Accuracy-critical ones are underpaying for verification.
Why one number can't be your procurement criterion
Here is the tell. FrontierCode found Opus 5 scoring better at medium reasoning effort than at high effort — a non-monotonic result that shouldn't happen if the aggregate index tracked real capability cleanly. Practitioners simultaneously report the model feels dramatically better than its +1 ECI gain over Opus 4.8 suggests. The index is diverging from lived experience in both directions at once, which is the strongest argument yet for owning your own eval set. The leaderboard tells you who to shortlist, not what to ship.
Leaderboard rank buys capability, not correctness — those are separate purchases, and only one of them is on your risk register.
What to do
Re-run unit economics this sprint on every AI feature shelved for margin in the last two quarters, and flag which ones clear the bar at Opus 5's pricing before the next prioritization review
Add a factual-reliability column to your model-selection matrix and re-score each AI surface by task risk this sprint, defaulting accuracy-critical flows away from the highest-benchmark model
Run a two-week spike swapping one output-heavy, non-accuracy-critical workload to GLM 5.2 and compare quality and cost against your incumbent on your own prompts