Sonnet 5.5 beats your ticket routing model by 11.9 points at 41% lower cost
New release·Sonnet 5.5 shipped Oct 4, 09:00·Scored 412 cases by 10:02
Your ticket routing runs GPT-5.6 Sol at $2.90 per task and scores 35.8%
Sonnet 5.5 scores 47.7% at $1.70 and now holds 4 of the 8 points on your frontier
- Quality
- 47.7%▲ +11.9 pts
- Cost per task
- $1.70▼ 0.6×
- Estimated monthly cost at 60,000 tasks
- $102k▼ −$72k
What changed
All runs
26 models · 412 cases · avg@3| Model | Quality | Cost / task | p95 | Tokens / task | Turns | Frontier |
|---|---|---|---|---|---|---|
Opus 5.5 Max | 57.8% | $13.43 | 190.8s | 218,363 | 185 | Yes |
Opus 5.5 Extra High | 56.0% | $6.98 | 110.5s | 101,083 | 109 | No |
Opus 5.5 High | 56.0% | $3.97 | 72.5s | 58,972 | 68 | Yes |
Sonnet 5.5 Max | 55.4% | $9.60 | 121.7s | 171,204 | 142 | No |
Sonnet 5.5 High | 53.1% | $3.90 | 71.7s | 72,418 | 81 | Yes |
Opus 5.5 Medium | 52.5% | $2.90 | 52.3s | 41,560 | 52 | Yes |
Sonnet 5.5 Medium | 47.7% | $1.70 | 36.3s | 30,118 | 44 | Yes |
GPT-5.6 Sol Medium | 35.8% | $2.90 | 57.2s | 46,902 | 61 | No |
Where each model lands
Top right is best on both axes, faded dots are off the frontier
Speed against cost
Latency p95 · cost per taskSonnet 5.5 Minimal is the cheapest and fastest run, at $0.50 and 15.7s per task.
It matches production's 35.8% at a sixth of the cost. Sonnet 5.5 Medium is the pick, scoring 47.7% at $1.70 and 36.3s, both under production's $2.90 and 57.2s.
Quality against speed
Quality · latency p95Opus 5.5 High gives the best quality without the wait, 56.0% at a 72.5s p95.
Opus 5.5 Max adds only 1.8 points and takes 118.3s longer. Sonnet 5.5 Medium is twice as fast at 36.3s for 8.3 points less.
Token appetite
Quality · tokens per taskOpus 5.5 High is the most token-efficient top scorer, 56.0% on 58,972 tokens per task.
Gemini 3.8 Flash High burns 187,621 tokens for 39.7%, over three times as many for 16.3 points less. Sonnet 5.5 Medium uses 30,118, about a third fewer than production's 46,902.
Turns against quality
Quality · turns per taskOpus 5.5 Medium needs the fewest turns of any run above 50%, scoring 52.5% in 52 turns.
Sonnet 5.5 Medium takes 44 turns, 17 fewer than production, and scores 11.9 points higher. Opus 5.5 Max needs 185 turns for 57.8%.
How this was scored
- Cases
- 412
- from your support ticket history
- Runs
- avg@3
- 3 runs per case, scores averaged
- Grading
- Rubric
- grader v14
- Next run
- Next release
- Anthropic and OpenAI releases watched
Release history
| Date | Release | Champion | Result |
|---|---|---|---|
| Oct 4 | Sonnet 5.5 | Sonnet 5.5 | Better on ticket routing |
| Sep 9 | GPT-5.6 Sol | GPT-5.6 Sol | Better on ticket routing |
| Aug 12 | Opus 5.5 | Opus 5.5 | Better on refund approvals and reply drafting |
| Jun 24 | Fable 5.1 | Opus 5.1 | Same on refund approvals at 4× the cost |
| May 6 | GPT-5.5 | Opus 5.1 | Slower on reply drafting |
| Feb 18 | GPT-5.4 | GPT-5.4 | Better on ticket routing |