Score coding models on your repo, not SWE-bench
Bring the PRs, bugs, and migrations your team already closed. Klu turns them into cases that pass or fail, reruns them when a model ships, and tells you which workflows should switch.
Would your reviewers merge it?
Klu weights each run the way your senior reviewers do. A cheaper model can't win by passing the easy cases and breaking a caller on the hard ones.
Every case passes or fails
A case is one change your team already shipped, plus the checks a reviewer would make. Mark a check required and missing it fails the case, however clean the rest of the patch looks.
Each check comes back with the evidence the grader quoted from the patch, so you can see why a model lost the case.
Sonnet 5.5 leads terminal coding at $4
Terminal-Bench 4.0 gives a model a terminal and a real engineering task. Sonnet 5.5 finishes more tasks than Opus 5.5 or GPT-6 Astra. On ITBench incidents, Opus 5.5 ranks 12th of 21. Klu reruns these models on your repo.
Where engineering teams use Klu
Each one starts with a default way to score it. Change it if yours differs.
PR review
Review bots that comment on every pull request. Scored on the bugs your reviewers caught, with each comment they would dismiss counted against it.
Bug fixes
Issue in, patch out. Each case is a bug your team already fixed, scored against the merged fix. A patch that breaks another caller fails.
Test generation
Unit tests for new and changed code. A test counts when it targets the behavior that broke, not when it adds coverage or asserts on internals.
Migrations
Framework upgrades, deprecated APIs, language ports. Each file is scored against your hand-migrated version, and any behavior change fails it.
One line per workflow, scored on your repo
Only the new model reruns, on the same frozen cases. You see which workflows should switch, which stay, and the cases each model won and lost.
Find your best model
Autopilot for peak performance, lowest price
Checked when labs ship