Know which model answers from your docs, not around them

Klu reruns your Q&A evals when a new model ships. It checks that each answer cites the right doc, uses the current version, and says "not covered" when your docs have nothing.

Employee helpdeskGemini 4 Argon release
Switch to Gemini 4 Argon. It guessed on 2 of 80 unanswerable questions. Fable 5.1 guessed on 13.
Correct answers, Fable 5.189.2%→ today
Correct answers, Gemini 4 Argon87.5%▼ −1.7 pts
Guessed when docs had no answer, Fable 5.116.3%▲ 6.5×
Cost per answer, Gemini 4 Argon$0.006▼ −81%
Scored on 480 real employee questions. 80 have no answer in your docs, and Gemini 4 Argon guessed on 2 of them.

A right answer from the wrong doc still fails

Most bad answers read fine. They quote last year's policy, or the US page to someone in London, or fill a gap in your docs with a guess. Klu scores each of those on its own line, and a guess on a question your docs don't cover scores zero.

01
Correct answerMatches the answer your team reviewed
35%
02
GroundedEvery claim is in the passages it got
25%
03
Right sourceCites the current doc that holds the answer
15%
04
Says when it doesn't knowDeclines and routes when docs are silent
15%
05
Cost per answerAt your monthly question volume
10%

Answers that read fine and still fail

Each case holds the question, the passages your retriever returned, and the answer your reviewers expect. Every model gets the same passages, so a change in score comes from the model, not from search.

Wrong sourceCase 112
Asked by an employee in the London officeHow much paid parental leave do I get?
Retrieved passages
Parental leave, USConfluence · PeopleCited
Parental leave, UKConfluence · People
Haiku 4.5

You get 16 weeks of paid parental leave.

Correct answerFail
GroundedPass
Right sourceFail

Quoted the US policy word for word. The UK policy applies to London staff.

Stale docCase 241
Asked in #ask-financeWhat's the meal per diem in New York?
Retrieved passages
Travel and expense policy 2024SharePoint · FinanceArchivedCited
Travel and expense policy 2026SharePoint · Finance
Sonnet 5.5

Meals in New York are covered up to $75 a day.

Correct answerFail
GroundedPass
Right sourceFail

Cited the archived policy. The 2026 policy raised New York to $95.

Not in the docsCriticalCase 377
Asked in #ask-peopleCan I take a sabbatical after five years?
Retrieved passages
Paid time offHandbook · Time off
Leave of absenceHandbook · Time off
Gemini 4 Argon

The handbook doesn't mention a sabbatical program. The People team can tell you whether one exists.

Says when it doesn't knowPass
GroundedPass
Correct answerPass

Fable 5.1 described a four-week sabbatical here. Critical case, so it scored zero.

The most accurate models guess the most

Omniscience asks closed-book questions and scores two things: right answers, and declining instead of inventing. The most accurate models invent the most. Gemini 4 Argon declines far more often than any other flagship. Klu reruns these models on your docs.

Omniscience179 models · Oct 7, 2026
AccuracyDeclines instead of guessing
1
Fable 5.1 Max$20 per 1M
67.2%
27%
2
Opus 5.5 Max$8 per 1M
66.2%
41%
3
GPT-6 Astra Max$20 per 1M
62.6%
49%
4
GPT-6.1 Sol Max$4 per 1M
62.1%
46%
7
Sonnet 5.5 Max$4 per 1M
53.9%
53%
9
Gemini 4 Argon High$4 per 1M
49.9%
85%
15
DeepSeek V4.1 Flash Max$0.52 per 1M
46.4%4%

Where knowledge teams use Klu

Each one starts with a default way to score it. Change it if yours differs.

Right source, right answer

Employee helpdesk

HR, IT, and finance questions in Slack or Teams, answered from your handbook, Confluence, and SharePoint. Scored on the answer and the policy it cites.

Every claim cited

Help center answers

Customer questions answered from your help articles and product docs. Each claim must trace to an article, and a retired plan or feature counts as wrong.

Current version only

Runbooks and engineering docs

On-call and developer questions over runbooks, API references, and design docs. Steps from a retired runbook fail the case.

Holds up over follow-ups

Multi-turn copilot

Simulated users who narrow, correct, or change the question mid-chat. Scored across the whole conversation, not only the first answer.

One verdict per assistant, with the reason

When a model ships, Klu reruns each eval set and tells you which assistants should switch, which should stay, and which failure decided it.

Gemini 4 Argon releaseReport ready 4 hrs after release
Help centerSwitchGemini 4 Argon
▲ +2.9 pts−52% cost per answer
Employee helpdeskSwitchGemini 4 Argon
▼ 6.5× fewer guesseson questions docs skip
RunbooksKeepOpus 5.5
▼ −4.1 ptscited retired runbooks
Copilot chatKeepHaiku 4.5
→ +0.8 ptssame score, 3.1× the cost

Find your best model

Autopilot for peak performance, lowest price
Checked when labs ship