Know a model gets the science right before you switch

Klu reruns your R&D evals the day a new model ships and shows whether it cites, calculates, and drafts protocols the way your scientists would.

Literature synthesisSonnet 5.5 release
Keep Opus 5.5 High. Sonnet 5.5 miscites a paper in 3.2× as many answers.
Answers with a miscitation, Opus 5.5 High2.5%▼ best
Answers with a miscitation, Sonnet 5.58.0%▲ 3.2×
Cost per answer, Opus 5.5 High$1.36→ today
Cost per answer, Sonnet 5.5$0.58▼ −57%
Scored on 200 review questions against your scientists' reference answers and the papers each answer cites.

Scored the way a PI reads a draft

Klu weights each run the way your reviewers mark up a draft, so a cheaper model can't win by citing a paper that says something else or by slipping from nM to µM.

01
Scientific accuracyAgrees with your scientists' reference
35%
02
Citation supportThe cited source says what the claim says
25%
03
Numbers and unitsValues, units, and sequences exact
15%
04
Calibrated claimsSpecies, n, and certainty as reported
15%
05
Cost per taskAt your monthly volume
10%

GPT-6 and Claude lead science. Nothing else clears 14%

Terminal-Bench-Science gives a model a terminal and a science task. GPT-6 and Claude hold the top five spots, and no other model clears 14%. On Humanity's Last Exam, Opus 5.5 is first. Klu reruns these models on your assays and protocols.

Terminal-Bench-Science21 models · Oct 7, 2026
Tasks passed
1
GPT-6 Astra Max$20 per 1M
63.3%
2
Opus 5.5 Xhigh$8 per 1M
61.9%
3
GPT-6.1 Sol Max$4 per 1M
58.1%
4
Sonnet 5.5 Max$4 per 1M
53.3%
5
Fable 5.1 Max$20 per 1M
43.3%
10
GLM-5.3 Max$2.15 per 1M
9.5%
11
DeepSeek V4.1 Flash Max$0.52 per 1M
9.0%
Humanity's Last ExamQuestions right185 models
1
Opus 5.5 Max$8 per 1M
61.4%
2
Fable 5.1 Max$20 per 1M
59.1%
3
Gemini 4 Argon High$4 per 1M
57.1%

Where research teams use Klu

Biology first. Each one starts with a default way to score it, and the same rubrics carry over to chemistry and materials work.

Every claim cited

Literature synthesis

Answer a research question from a set of papers, preprints, patents, and internal reports. A citation the source does not support fails the answer, and so does a mouse result stated as human.

Checked against known answers

Protein and molecule tasks

Annotate domains, predict variant effects, design primers, reason over SMILES. Scored against your assay results and curated annotations, not the model's confidence.

Unsafe step fails

Protocol drafting

Turn a methods section into a bench protocol for your instruments. Volumes and concentrations must match the reference, and a missing safety step fails the protocol.

Field-level accuracy

Assay data extraction

Pull IC50, Kd, n, and conditions from papers, patents, and supplementary tables. A value in µM when the source says nM counts as wrong.

Open any answer and see why it failed

Every criterion score quotes the passage it rests on, so a scientist can check the judge's call against the source in one read.

Mark citation support as required and a single miscitation scores the whole answer zero, however well the rest reads.

Lit synthesis · Case 112 · Sonnet 5.5Score 0
QuestionWhich compounds in the attached set slowed tumor growth in vivo, and by how much?
AnswerCompound 14 cut tumor volume by 62% in patients with KRAS G12C tumors [S3]. Compound 9 slowed growth by 35% [S5].
Citation support25% · required0.0Failed
S3, Results: "Tumor volume fell 62% in PDX-bearing mice (n = 8) by day 21." No patient data in S3.
Calibrated claims15%0.2Failed
States a mouse xenograft result as a patient result.
Numbers and units15%1.0Passed
62% and day 21 match S3, Table 2.
Scientific accuracy35%0.8Passed
Names the same two compounds as the reference answer.

One line per workflow, scored on the science

Every release gets the same treatment: which workflows should switch, which stay, and why, in terms your head of R&D can sign off on.

Sonnet 5.5 releaseReport ready 1 hr after release
Lit synthesisKeepOpus 5.5 High
▼ 3.2× miscitesSonnet 5.5's best
Assay extractionSwitchSonnet 5.5 Medium
▲ +2.8 pts−46% cost per paper
Variant effectsSwitchSonnet 5.5 Medium
▲ +5.1 ptsbeats GPT-5.6 Sol
Protocol draftsKeepOpus 5.5 High
→ +0.4 ptssame within 2 pts

Find your best model

Autopilot for peak performance, lowest price
Checked when labs ship