Switch to a cheaper model without resolving fewer tickets

Klu runs your past tickets through each model you're weighing, scores them on your QA scorecard, and puts a price on every one. When a new model ships, it reruns them that day.

Support chat agentMuse Spark 1.3 rerun
Switch chat to Muse Spark 1.3. It resolves as many chats as Opus 5.5 High for 46% less.
Resolved correctly, Opus 5.5 High68.3%→ today
Resolved correctly, Muse Spark 1.369.0%▲ +0.7 pts
Cost per chat, Opus 5.5 High$0.41→ today
Cost per chat, Muse Spark 1.3$0.22▼ −46%
Scored on 400 of your past chats, with Klu playing the customer. Neither model refunded outside policy.

Resolved and on policy first. Then cost per ticket.

Klu weights each ticket the way your QA leads do. A refund over the limit or a missed escalation fails the ticket outright, so a cheaper model can't win by handing out credits.

01
ResolutionSolved what the customer asked
35%
02
Policy complianceRefunds, credits, promises within limits
25%
03
EscalationHands off when it should, with context
15%
04
Tone and clarityYour voice, one clear next step
10%
05
Cost per ticketAt your monthly volume
15%

Models under $3 lead banking support

τ³-Bench Banking scores multi-turn chats with a simulated bank customer. A chat passes only if it ends within policy. The top five all cost $3 or less, and the $20 Fable 5.1 is seventh. Klu reruns these models on your tickets.

τ³-Bench Banking81 models · Oct 7, 2026
Conversations passed
1
Muse Spark 1.3 Max$2 per 1M
50.5%
2
GLM-5.3 Max$2.15 per 1M
50.3%
3
Qwen3.8 2.4T$3 per 1M
49.1%
4
Qwen3.8 27B Xhigh$1.13 per 1M
48.0%
5
Qwen3.8 Max$3 per 1M
47.8%
7
Fable 5.1 Max$20 per 1M
47.2%
11
GPT-6 Astra Xhigh$20 per 1M
43.1%

Where support teams use Klu

Each one starts with a default way to score it. Change it if yours differs.

Resolved, on policy

AI agent conversations

The bot on your chat widget or email queue. Klu plays the customer for up to 12 turns, then checks the outcome, any refund, and when it handed off.

Matches what your agent sent

Reply drafts

Suggested replies your agents edit in Zendesk, Intercom, or Front. Scored against the reply that went out, plus policy and tone.

Exact queue match

Triage and routing

Queue, priority, and language on every inbound ticket. Exact match against where your team routed it, including tickets that raise three issues at once.

Agreement with your leads

QA scoring

Auto-QA that grades every conversation on your scorecard. Scored on how often it gives the marks your QA leads gave.

Switch or keep, one line per workflow

Each release gets the same report: which workflows should switch, which stay, the score change, and what it does to cost per ticket.

Muse Spark 1.3 rerunReport ready 1 hr after release
Support chat agentKeepMuse Spark 1.3
▼ −0.7 ptsSonnet 5.5 at 1.9× the cost
Reply draftsSwitchSonnet 5.5
▲ +3.4 pts−31% cost per draft
Ticket triageKeepHaiku 4.5
→ 0.0 ptsSonnet 5.5 at 3× the cost
QA scoringKeepOpus 5.5 High
▼ −6.1 pts agreementSonnet 5.5's best

Read the chats it lost

Muse Spark 1.3 ties Opus 5.5 High on resolution, and a tie still hides losses. Klu lists every chat where the new model scored lower, names the criterion it missed, and puts both replies side by side.

Support chat agent · Muse Spark 1.3 vs Opus 5.5 HighLost on 11 of 400 chats
Chargeback threat after a downgradeEscalation
Argued the downgrade date before handing off to a team lead
▼ 58 pts
One order, two packagesResolution
Tracked the first package and closed the chat
▼ 47 pts
Address change after shippingResolution
Told the customer to wait instead of starting a redirect
▼ 39 pts
Duplicate charge in MarchTone and clarity
Refunded the right charge but never said when it lands
▼ 22 pts
and 7 moreOpen both replies

Find your best model

Autopilot for peak performance, lowest price
Checked when labs ship