Score coding models on your repo, not SWE-bench

Bring the PRs, bugs, and migrations your team already closed. Klu turns them into cases that pass or fail, reruns them when a model ships, and tells you which workflows should switch.

Bug fixesSonnet 5.5 release
Switch bug fixes to Sonnet 5.5. It fixes 11 more of 240 bugs at 47% less.
Pass rate, Sonnet 5.581.3%▲ best
Pass rate, Opus 5.5 High76.7%▼ −4.6 pts
Cost per fix, Opus 5.5 High$1.84→ today
Cost per fix, Sonnet 5.5$0.97▼ −47%
240 bugs your team fixed this year, each scored against the merged fix.

Would your reviewers merge it?

Klu weights each run the way your senior reviewers do. A cheaper model can't win by passing the easy cases and breaking a caller on the hard ones.

01
Correct fixSame behavior as the merged change
40%
02
No regressionsCallers, schemas, auth, data
25%
03
Fits the codebaseYour patterns, smallest diff, no new deps
15%
04
Signal over noiseNo invented bugs or busywork
10%
05
Cost per taskAt your monthly PR volume
10%

Every case passes or fails

A case is one change your team already shipped, plus the checks a reviewer would make. Mark a check required and missing it fails the case, however clean the rest of the patch looks.

Each check comes back with the evidence the grader quoted from the patch, so you can see why a model lost the case.

Bug fixes · case 37 of 240Merged fix as reference
IssueWebhook retries ignore Retry-After. After a 429, deliveries retry within a second and get throttled again.
webhooks/deliver.py
if resp.status == 429:
- delay = backoff(attempt)
+ delay = retry_after(resp) or backoff(attempt)
ChecksSonnet 5.5Opus 5.5 High
Reads Retry-After before retryingRequired
Leaves backoff() unchanged for its other callersRequired
Adds a test that fails on the old code
Keeps the diff inside webhooks/
Case resultPassFail
Opus 5.5 High rewrote backoff() in lib/retry.py, which 14 other callers use. A required check failed, so the case scores zero.

Sonnet 5.5 leads terminal coding at $4

Terminal-Bench 4.0 gives a model a terminal and a real engineering task. Sonnet 5.5 finishes more tasks than Opus 5.5 or GPT-6 Astra. On ITBench incidents, Opus 5.5 ranks 12th of 21. Klu reruns these models on your repo.

Terminal-Bench 4.085 models · Oct 7, 2026
Tasks passed
1
Sonnet 5.5 Max$4 per 1M
63.6%
2
GPT-6 Astra Xhigh$20 per 1M
59.6%
3
Opus 5.5 Max$8 per 1M
59.6%
4
Gemini 4 Argon High$4 per 1M
57.1%
5
GPT-6.1 Sol Max$4 per 1M
56.1%
6
Fable 5.1 Xhigh$20 per 1M
55.1%
ITBenchIncidents resolved21 models
1
Step 5 Preview$1.43 per 1M
55.6%
2
Gemini 3.8 Flash High$1.50 per 1M
52.5%
3
GLM 5.3 Flash$0.24 per 1M
51.2%
12
Opus 5.5 Max$8 per 1M
38.2%

Where engineering teams use Klu

Each one starts with a default way to score it. Change it if yours differs.

Real bugs, few false alarms

PR review

Review bots that comment on every pull request. Scored on the bugs your reviewers caught, with each comment they would dismiss counted against it.

Root cause, not symptom

Bug fixes

Issue in, patch out. Each case is a bug your team already fixed, scored against the merged fix. A patch that breaks another caller fails.

Fails on the old code

Test generation

Unit tests for new and changed code. A test counts when it targets the behavior that broke, not when it adds coverage or asserts on internals.

Behavior unchanged

Migrations

Framework upgrades, deprecated APIs, language ports. Each file is scored against your hand-migrated version, and any behavior change fails it.

One line per workflow, scored on your repo

Only the new model reruns, on the same frozen cases. You see which workflows should switch, which stay, and the cases each model won and lost.

Sonnet 5.5 releaseReport ready 1 hr after release
PR reviewSwitchSonnet 5.5 Medium
▲ +3.1 pts−44% cost per review
MigrationsSwitchSonnet 5.5 High
▲ +1.8 pts−31% cost per file
Bug fixesSwitchSonnet 5.5
▲ +4.6 pts−47% cost per fix
Test generationKeepHaiku 4.5
→ +0.4 ptsnot worth 3× the cost

Find your best model

Autopilot for peak performance, lowest price
Checked when labs ship