Know a cheaper model gets every field right before you switch
Klu scores each model against the values your team already keyed, on your own invoices, claims, policies, and leases. When a new model ships, Klu reruns the set and tells you which document types can move.
A blank goes to review. A wrong value gets paid.
Klu weights each run the way your QA team would. A blank field lands in the exception queue. A value that isn't on the page posts straight through, so it costs more. Output that fails your schema fails the document.
See which field each model got wrong
Open any document in the run and compare each model's output with the values your team keyed. A swapped day and month, a PO number that isn't on the page, and two line items merged into one each count as their own miss.
Haiku 4.5 costs less per invoice. On this vendor's layout it also posts a PO number nobody issued.
| Field | Keyed | GPT-6.1 Sol | Haiku 4.5 |
|---|---|---|---|
| Invoice number | INV-20931 | INV-20931 | INV-20931 |
| Invoice date | 2026-03-04 | 2026-03-04 | 2026-04-03day and month swapped |
| PO number | null | null | PO-55120not on the page |
| Line items | 14 rows | 14 rows | 13 rowsrows 6 and 7 merged |
| Subtotal | 11,840.00 | 11,840.00 | 11,840.00 |
| Tax | 947.20 | 947.20 | 947.20 |
| Total | 12,787.20 | 12,787.20 | 12,787.20 |
| Fields right | 7 | 7 of 7 | 4 of 7 |
The best model passes 32% of real PDF tasks
GDP.pdf asks real questions about 100 real PDFs, such as leases, insurance policies, and datasheets. A task passes only if every criterion does. OpenAI holds the top two spots. Klu reruns these models on your documents, field by field.
Where extraction teams use Klu
Each one starts with a default way to score it. Change it if yours differs.
Invoices and receipts
Header fields, line items, tax, and totals into your AP schema. Scored per field against keyed values, with totals checked against the lines.
Claims intake
Loss notices, adjuster notes, repair estimates, and medical bills into the claim record. Policy number, date of loss, and amounts must match exactly.
Policies and loss runs
ACORD applications, dec pages, and carrier loss runs into underwriting fields. Five years of losses in each carrier's format, scored row by row.
Leases and contracts
Rent schedules, escalations, renewal options, and notice windows. A wrong notice date costs more than a blank one, so it scores that way.
One line per document type, scored field by field
Every release gets the same rerun on the same documents. You see which document types should move, which should stay, and the fields behind each call. Klu never switches a model for you.
Find your best model
Autopilot for peak performance, lowest price
Checked when labs ship