Know which model answers from your docs, not around them
Klu reruns your Q&A evals when a new model ships. It checks that each answer cites the right doc, uses the current version, and says "not covered" when your docs have nothing.
A right answer from the wrong doc still fails
Most bad answers read fine. They quote last year's policy, or the US page to someone in London, or fill a gap in your docs with a guess. Klu scores each of those on its own line, and a guess on a question your docs don't cover scores zero.
Answers that read fine and still fail
Each case holds the question, the passages your retriever returned, and the answer your reviewers expect. Every model gets the same passages, so a change in score comes from the model, not from search.
You get 16 weeks of paid parental leave.
Quoted the US policy word for word. The UK policy applies to London staff.
Meals in New York are covered up to $75 a day.
Cited the archived policy. The 2026 policy raised New York to $95.
The handbook doesn't mention a sabbatical program. The People team can tell you whether one exists.
Fable 5.1 described a four-week sabbatical here. Critical case, so it scored zero.
The most accurate models guess the most
Omniscience asks closed-book questions and scores two things: right answers, and declining instead of inventing. The most accurate models invent the most. Gemini 4 Argon declines far more often than any other flagship. Klu reruns these models on your docs.
Where knowledge teams use Klu
Each one starts with a default way to score it. Change it if yours differs.
Employee helpdesk
HR, IT, and finance questions in Slack or Teams, answered from your handbook, Confluence, and SharePoint. Scored on the answer and the policy it cites.
Help center answers
Customer questions answered from your help articles and product docs. Each claim must trace to an article, and a retired plan or feature counts as wrong.
Runbooks and engineering docs
On-call and developer questions over runbooks, API references, and design docs. Steps from a retired runbook fail the case.
Multi-turn copilot
Simulated users who narrow, correct, or change the question mid-chat. Scored across the whole conversation, not only the first answer.
One verdict per assistant, with the reason
When a model ships, Klu reruns each eval set and tells you which assistants should switch, which should stay, and which failure decided it.
Find your best model
Autopilot for peak performance, lowest price
Checked when labs ship