Switch to a cheaper model without resolving fewer tickets
Klu runs your past tickets through each model you're weighing, scores them on your QA scorecard, and puts a price on every one. When a new model ships, it reruns them that day.
Resolved and on policy first. Then cost per ticket.
Klu weights each ticket the way your QA leads do. A refund over the limit or a missed escalation fails the ticket outright, so a cheaper model can't win by handing out credits.
Models under $3 lead banking support
τ³-Bench Banking scores multi-turn chats with a simulated bank customer. A chat passes only if it ends within policy. The top five all cost $3 or less, and the $20 Fable 5.1 is seventh. Klu reruns these models on your tickets.
Where support teams use Klu
Each one starts with a default way to score it. Change it if yours differs.
AI agent conversations
The bot on your chat widget or email queue. Klu plays the customer for up to 12 turns, then checks the outcome, any refund, and when it handed off.
Reply drafts
Suggested replies your agents edit in Zendesk, Intercom, or Front. Scored against the reply that went out, plus policy and tone.
Triage and routing
Queue, priority, and language on every inbound ticket. Exact match against where your team routed it, including tickets that raise three issues at once.
QA scoring
Auto-QA that grades every conversation on your scorecard. Scored on how often it gives the marks your QA leads gave.
Switch or keep, one line per workflow
Each release gets the same report: which workflows should switch, which stay, the score change, and what it does to cost per ticket.
Read the chats it lost
Muse Spark 1.3 ties Opus 5.5 High on resolution, and a tie still hides losses. Klu lists every chat where the new model scored lower, names the criterion it missed, and puts both replies side by side.
Find your best model
Autopilot for peak performance, lowest price
Checked when labs ship