IndustriesPricingServicesBenchmarks
CyberTalk to salesStart
Start
Get started

Sonnet 5.5 beats your ticket routing model by 11.9 points at 41% lower cost

New release·Sonnet 5.5 shipped Oct 4, 09:00·Scored 412 cases by 10:02

Your ticket routing runs GPT-5.6 Sol at $2.90 per task and scores 35.8%
Sonnet 5.5 scores 47.7% at $1.70 and now holds 4 of the 8 points on your frontier

Quality
47.7%▲ +11.9 pts
Cost per task
$1.70▼ 0.6×
Estimated monthly cost at 60,000 tasks
$102k▼ −$72k
Fable 5.1GPT-5.6 SolGrok 4.7Opus 5.5 MaxSonnet 5.5ProductionNew champion50%40%30%$15$9$3
Your frontierTicket routingFrontierSonnet 5.5Opus 5.5Others

What changed

All runs

26 models · 412 cases · avg@3
ModelQualityCost / taskp95Tokens / taskTurnsFrontier
Opus 5.5 Max
57.8%
$13.43
190.8s
218,363
185
Yes
Opus 5.5 Extra High
56.0%
$6.98
110.5s
101,083
109
No
Opus 5.5 High
56.0%
$3.97
72.5s
58,972
68
Yes
Sonnet 5.5 Max
55.4%
$9.60
121.7s
171,204
142
No
Sonnet 5.5 High
53.1%
$3.90
71.7s
72,418
81
Yes
Opus 5.5 Medium
52.5%
$2.90
52.3s
41,560
52
Yes
Fable 5.1 Max
51.8%
$17.30
212.1s
164,065
171
No
Fable 5.1 Extra High
51.6%
$13.00
167.3s
121,347
133
No
Fable 5.1 High
49.2%
$9.10
117.9s
84,471
97
No
Sonnet 5.5 Medium
47.7%
$1.70
36.3s
30,118
44
Yes
Fable 5.1 Medium
46.8%
$7.10
98.3s
65,143
78
No
Grok 4.7 Max
46.3%
$6.00
90.2s
117,907
126
No
Fable 5.1 Low
45.1%
$5.50
79.4s
48,564
63
No
Grok 4.7 High
43.9%
$4.70
72.3s
91,727
101
No
Opus 5.5 Low
43.8%
$1.30
31.2s
20,049
31
Yes
GPT-5.6 Sol Extra High
41.8%
$8.30
138.7s
130,686
152
No
Grok 4.7 Medium
41.6%
$3.50
57.1s
70,380
79
No
Gemini 3.8 Flash High
39.7%
$4.70
74.3s
187,621
164
No
Sonnet 5.5 Low
39.2%
$0.70
20.3s
12,589
24
Yes
GPT-5.6 Sol High
37.8%
$4.40
84.8s
71,085
88
No
Gemini 3.8 Flash Medium
37.4%
$4.10
64.5s
162,074
139
No
Sonnet 5.5 Minimal
35.8%
$0.50
15.7s
8,572
19
Yes
GPT-5.6 Sol Medium
35.8%
$2.90
57.2s
46,902
61
No
Grok 4.7 Low
33.1%
$1.60
29.7s
30,930
41
No
GPT-5.6 Sol Low
31.0%
$1.80
37.8s
29,061
39
No
GPT-5.6 Sol Minimal
24.5%
$0.90
20s
14,078
22
No

Where each model lands

Top right is best on both axes, faded dots are off the frontier

Expensive and fastCheap and fastExpensive and slowCheap and slow$0$4.25$180s72s220sCost per task, cheaper to the rightLatency p95, faster upChampionProduction

Speed against cost

Latency p95 · cost per task

Sonnet 5.5 Minimal is the cheapest and fastest run, at $0.50 and 15.7s per task.

It matches production's 35.8% at a sixth of the cost. Sonnet 5.5 Medium is the pick, scoring 47.7% at $1.70 and 36.3s, both under production's $2.90 and 57.2s.

Slow and higher qualityFast and higher qualitySlow and lower qualityFast and lower quality0s72s220s20%44.5%60%Latency p95, faster to the rightQualityChampionProduction

Quality against speed

Quality · latency p95

Opus 5.5 High gives the best quality without the wait, 56.0% at a 72.5s p95.

Opus 5.5 Max adds only 1.8 points and takes 118.3s longer. Sonnet 5.5 Medium is twice as fast at 36.3s for 8.3 points less.

Many tokens and higher qualityFew tokens and higher qualityMany tokens and lower qualityFew tokens and lower quality071k230k20%44.5%60%Tokens per task, fewer to the rightQualityChampionProduction

Token appetite

Quality · tokens per task

Opus 5.5 High is the most token-efficient top scorer, 56.0% on 58,972 tokens per task.

Gemini 3.8 Flash High burns 187,621 tokens for 39.7%, over three times as many for 16.3 points less. Sonnet 5.5 Medium uses 30,118, about a third fewer than production's 46,902.

Many turns and higher qualityFew turns and higher qualityMany turns and lower qualityFew turns and lower quality08020020%44.5%60%Turns per task, fewer to the rightQualityChampionProduction

Turns against quality

Quality · turns per task

Opus 5.5 Medium needs the fewest turns of any run above 50%, scoring 52.5% in 52 turns.

Sonnet 5.5 Medium takes 44 turns, 17 fewer than production, and scores 11.9 points higher. Opus 5.5 Max needs 185 turns for 57.8%.

How this was scored

Cases
412
from your support ticket history
Runs
avg@3
3 runs per case, scores averaged
Grading
Rubric
grader v14
Next run
Next release
Anthropic and OpenAI releases watched

Release history

DateReleaseChampionResult
Oct 4Sonnet 5.5Sonnet 5.5Better on ticket routing
Sep 9GPT-5.6 SolGPT-5.6 SolBetter on ticket routing
Aug 12Opus 5.5Opus 5.5Better on refund approvals and reply drafting
Jun 24Fable 5.1Opus 5.1Same on refund approvals at 4× the cost
May 6GPT-5.5Opus 5.1Slower on reply drafting
Feb 18GPT-5.4GPT-5.4Better on ticket routing
Automatically updates with each model releaseRun this on your workflows →

Find your frontier with the most intelligent applied researcher

ProductStudioObserveEvaluateOptimizeIntegrations
GuidesLLM OpsLLM EvaluationEngineeringOptimize Apps
ResourcesLLM LeaderboardBenchmarksGlossaryBlog
© 2022–2026 Klu, Inc.
PrivacyTerms