See what a new model misses before it reaches a chart

Klu reruns your clinical evals within a day of every model release. You see, per workflow, whether the new model drops allergies, invents findings, or files urgent messages as routine, and what it costs.

Visit notesSonnet 5.5 rerun
Switch to Sonnet 5.5. It missed a critical finding in 5 of 600 notes. Opus 5.5 High missed 11.
Notes with a critical miss, Sonnet 5.50.8%▼ best
Notes with a critical miss, Opus 5.5 High1.8%▲ 2.2×
Cost per note, Opus 5.5 High$0.42→ today
Cost per note, Sonnet 5.5$0.19▼ −55%
Scored on 600 de-identified visits against the notes your physicians signed.

A missed allergy fails the note

The default rubric for ambient visit notes, built from how physician reviewers check a draft. Critical omissions and unsupported statements are required, so failing either scores the note zero. Cost goes on its own axis and can't buy back a miss.

01
Critical omissionsMeds, allergies, abnormal results, follow-up
30%
02
Unsupported statementsNothing the visit or chart doesn't back
30%
03
Attribution and negationWho said it, and what the patient denied
20%
04
Fits your templateRight sections, no padding
20%

Open a failed note and see the line it missed

The judge returns each criterion with the excerpt it used. Your clinicians calibrate it by picking the better of two notes, and can rescore any answer with a reason. Klu logs who did what.

Visit notes / Case 214 / Opus 5.5 HighJudge calibrated with 5 clinician picks
Urgent care visit, cellulitis of the left legPlan in the transcript: start cephalexin 500 mg four times daily
FailCritical omissionsRequired0.00
Transcript 03:12, patient: "Penicillin gives me hives." The note has no allergy section.
PassUnsupported statementsRequired1.00
Each finding traces to the transcript or the chart.
PassAttribution and negation0.90
"No fever or chills" stays a negative in the HPI.
PassFits your template0.85
SOAP sections in order, 240 words.
A required criterion failed, so the note scores 0.Sonnet 5.5 listed the allergy

Sonnet 5.5 beats Opus 5.5 on medical records

MLCR asks questions about long, fragmented medical records and grades the answers against experts. Sonnet 5.5 beats both pricier Claude models. OpenAI's flagships get about one record in three right. Klu reruns these models on your de-identified cases.

MLCR37 models · Oct 7, 2026
Cases right
1
Sonnet 5.5 Max$4 per 1M
75.0%
2
Fable 5.1 Max$20 per 1M
71.1%
3
Opus 5.5 Max$8 per 1M
66.7%
4
GLM 5.3 Flash$0.24 per 1M
51.1%
5
GLM-5.3 Max$2.15 per 1M
48.3%
8
GPT-6 Astra Max$20 per 1M
35.0%
9
GPT-6.1 Sol Max$4 per 1M
33.9%

Where healthcare teams use Klu

Each one starts with a default rubric and a list of misses that fail a case. Change either to match how your reviewers work.

Clinician rubric

Visit notes

Ambient visit notes and discharge summaries, scored against the note your physician signed. A missed allergy or an invented finding fails the note.

Criteria met, with evidence

Prior authorization

Match chart evidence to each payer criterion, draft the request, and flag what's missing. Scored against your utilization review nurses. A criterion marked met without evidence fails the case.

Code-level match

Medical coding

ICD-10-CM, CPT, and HCC codes from the encounter note, scored code by code against your coders' final codes. A code the note doesn't support fails the case.

Misses cost more

Patient messages

Triage portal messages and draft replies for your care team, scored against your triage nurses. An urgent message routed as routine fails, however good the draft.

One line per workflow, scored on what your reviewers check

Klu reruns each workflow on the new model and recommends switch or keep. A switch never happens on its own. Someone on your team picks the model.

Sonnet 5.5 releaseReport ready 1 hr after release
Visit notesSwitchSonnet 5.5
▼ 2.2× fewer critical misses5 of 600 notes, vs 11
Prior authSwitchSonnet 5.5 Medium
▲ +4.2 pts−38% cost per case
Medical codingKeepOpus 5.5 Medium
→ +0.6 ptsinside the ±2 pt band
Patient messagesSwitchSonnet 5.5 Medium
▲ +2.6 ptsno urgent misses, −46% cost

Find your best model

Autopilot for peak performance, lowest price
Checked when labs ship