Know a cheaper model gets every field right before you switch

Klu scores each model against the values your team already keyed, on your own invoices, claims, policies, and leases. When a new model ships, Klu reruns the set and tells you which document types can move.

Invoice extractionGPT-6.1 Sol release
Switch invoices to GPT-6.1 Sol. It matches Opus 5.5 High field for field at 45% less.
Field accuracy, GPT-6.1 Sol99.1%▲ +0.2 pts
Every field right, GPT-6.1 Sol93.8%▲ +0.9 pts
Per 1k invoices, Opus 5.5 High$39.70→ today
Per 1k invoices, GPT-6.1 Sol$21.80▼ −45%
Scored on 800 invoices from 140 vendors against the values your AP team keyed.

A blank goes to review. A wrong value gets paid.

Klu weights each run the way your QA team would. A blank field lands in the exception queue. A value that isn't on the page posts straight through, so it costs more. Output that fails your schema fails the document.

01
Field accuracyMatches the keyed value after normalizing
40%
02
No invented valuesBlank when the page is blank
25%
03
Tables and line itemsEvery row, and totals that reconcile
20%
04
Cost per documentAt your monthly volume
15%

See which field each model got wrong

Open any document in the run and compare each model's output with the values your team keyed. A swapped day and month, a PO number that isn't on the page, and two line items merged into one each count as their own miss.

Haiku 4.5 costs less per invoice. On this vendor's layout it also posts a PO number nobody issued.

INV-20931 · 2 pagesDocument 214 of 800
FieldKeyedGPT-6.1 SolHaiku 4.5
Invoice numberINV-20931INV-20931INV-20931
Invoice date2026-03-042026-03-042026-04-03day and month swapped
PO numbernullnullPO-55120not on the page
Line items14 rows14 rows13 rowsrows 6 and 7 merged
Subtotal11,840.0011,840.0011,840.00
Tax947.20947.20947.20
Total12,787.2012,787.2012,787.20
Fields right77 of 74 of 7

The best model passes 32% of real PDF tasks

GDP.pdf asks real questions about 100 real PDFs, such as leases, insurance policies, and datasheets. A task passes only if every criterion does. OpenAI holds the top two spots. Klu reruns these models on your documents, field by field.

GDP.pdf85 models · Oct 7, 2026
Tasks fully passed
1
GPT-6 Astra Xhigh$20 per 1M
32.2%
2
GPT-6.1 Sol High$4 per 1M
32.0%
3
Opus 5.5 High$8 per 1M
28.8%
4
Fable 5.1 Low$20 per 1M
28.0%
5
Muse Spark 1.3 Max$2 per 1M
26.6%
6
Sonnet 5.5 Max$4 per 1M
25.8%
13
Gemini 4 Argon High$4 per 1M
21.8%

Where extraction teams use Klu

Each one starts with a default way to score it. Change it if yours differs.

Field and line-item match

Invoices and receipts

Header fields, line items, tax, and totals into your AP schema. Scored per field against keyed values, with totals checked against the lines.

Field-level accuracy

Claims intake

Loss notices, adjuster notes, repair estimates, and medical bills into the claim record. Policy number, date of loss, and amounts must match exactly.

Every row, every year

Policies and loss runs

ACORD applications, dec pages, and carrier loss runs into underwriting fields. Five years of losses in each carrier's format, scored row by row.

Dates and dollars

Leases and contracts

Rent schedules, escalations, renewal options, and notice windows. A wrong notice date costs more than a blank one, so it scores that way.

One line per document type, scored field by field

Every release gets the same rerun on the same documents. You see which document types should move, which should stay, and the fields behind each call. Klu never switches a model for you.

GPT-6.1 Sol releaseReport ready 1 hr after release
InvoicesSwitchGPT-6.1 Sol
▲ +0.2 pts−45% cost per invoice
Claims intakeKeepOpus 5.5 High
▼ −1.8 ptsmisreads handwritten amounts
Loss runsKeepOpus 5.5 High
→ 0.0 ptsnothing beat it
LeasesSwitchGPT-6.1 Sol High
▲ +2.6 ptsfewer wrong notice dates

Find your best model

Autopilot for peak performance, lowest price
Checked when labs ship