Sol and Sonnet 5.5 Make Per-Token Price the Wrong Unit
Two cheaper models shipped in one week, and their savings sit in token counts, effort settings and failure costs that list prices never show.
The best argument for Sol is that independent evaluators agree with each other. AINews collected three per-task cost ratios against Astra, and they cluster tightly: ~4.5x from Artificial Analysis, ~4.5x implied by Roboflow's 78% saving, and ~5x from a 105-bug planted-bug run. The quality gap is the weak part. In the planted-bug test Sol found 44 bugs and Astra found 45. The two-proportion standard error is about 6.8 points, so z ≈ 0.14. That is a statistical tie, not a small loss. Roboflow's 81.6 vs 83.6 mAP@50 arrives without a dataset size or a CI. AA's one-point index gap is contested by Theo's Codex-harness runs, which scored much higher.
Same list price, different token counts
Sol and Sonnet 5.5 both list at $2 input / $10 output per million tokens. Any gap between them comes from tokens per task and success rate, and on tokens the two models move in opposite directions:
- Sol: AA measures it emitting 10–30% more output tokens than GPT-6 Sol. There is no 'none' reasoning effort, so every call pays a reasoning floor.
- Sonnet 5.5: Vals found it terser than its predecessor in 100% of paired tasks. Anthropic claims that at Low or Medium effort it beats Sonnet 5's best scores for about a tenth of the cost.
Output tokens cost 5x as much as input tokens. That makes effort level the biggest cost lever in any router, and max-effort Sonnet 5 configs are probably Pareto-dominated.
Sonnet 5.5 against Opus 5.5 is also a tie on current evidence. Sonnet leads on Terminal-Bench 4.0 (70.6% vs 66.4%) and trails on CursorBench 4.0 (55.5% vs 57.8%). Neither benchmark discloses its task count. At n=100, the standard error on the Terminal-Bench gap would be 6.6 points. The thing these scores don't tell you is that Sonnet 5.5 silently routes high-risk security requests back to Sonnet 5. Traffic labeled '5.5' is actually a mix of two models.
| Model | List price (per M) | Token behavior | Independent evidence | Trap |
|---|---|---|---|---|
| GPT-6.1 Sol | $2 / $10, $0.10 cached | +10–30% output, no 'none' effort | Four evaluators, all ties | Harness dispute on AA |
| Claude Sonnet 5.5 | $2 / $10 | Terser; effort is the dial | Vendor benchmarks only | Cyber reroute to Sonnet 5 |
| GPT-6 Astra | ~5x Sol | Baseline | Reference arm | 6.1 successor withheld |
Score cost per successful task
Turing Post's formula is (token cost + (1 − s) × failure-handling cost) / s, where s is the success rate. Its illustrative numbers:
- Astra: $0.50 of tokens per task at 85% success.
- Sol: $0.10 of tokens per task at 75% success.
- Both: $5 of cleanup per failed run.
That works out to $1.47 per success for Astra and $1.80 for Sol. When failures are cheap and retries idempotent, the ranking reverses. With the example's 5x gap in token cost per task, Sol then loses only if its success rate falls below a fifth of Astra's. Routing makes the same point. Fireworks' FireRouter cut coding-session cost from $15.36 to $6.63 while keeping 98.1% of Opus-only accuracy. Its A/B had no Sonnet 5.5-only arm, and that arm might capture most of the saving with no routing at all. Switching models mid-session also discards prompt-cache hits.
The smart move
Run a paired, pre-registered non-inferiority test on shadow traffic, with the margin fixed before anyone looks at results. An unpaired design needs about 6,500 sessions per arm to resolve a 2-point margin near 70% success. Replaying identical tasks through every arm and applying McNemar's test cuts that substantially. The readout means nothing until the judge and the sandbox are fixed, which the next section covers.
What to do
Shadow-route 5–10% of GPT-6 Astra traffic to GPT-6.1 Sol for two weeks starting this sprint. Pre-register a paired non-inferiority margin, and log output tokens, cache hit rate and cost per successful task.
Re-sweep Sonnet 5.5 at Low, Medium and High effort against Sonnet 5 and Opus 5.5 on your golden set this sprint. Retire every config that is Pareto-dominated.
Add a security-adjacent eval slice (auth, IAM, log parsing) before cutover. Track its token and latency distributions to detect silent reroutes to Sonnet 5.