Switch models when the numbers match, not when the SQL looks right

Klu reruns your text-to-SQL and transformation evals the day a new model ships and checks every result against your analysts' reference queries.

Text-to-SQL, finance questions
Switch to Sonnet 5.5 Medium. Result match holds at 58% lower cost per question.
Result match, Sonnet 5.5 Medium93.3%▲ +1.2 pts
Result match, Opus 5.5 High92.1%→ today
Silent wrong answers, Sonnet 5.50.6%▼ −0.2 pts
Cost per question, Sonnet 5.5$0.034▼ −58%
Scored on 480 questions against the results of your analysts' reference queries.

The rows have to match. The SQL can differ.

Two correct queries rarely look alike, so Klu scores what they return. Your rubric sets what counts as the same result, like row order, rounding, and column names. A confident wrong total costs the most.

01
Result matchSame rows and values as your reference
50%
02
Silent wrong answersA wrong number with no caveat
25%
03
Assumptions statedFiscal or calendar, gross or net
15%
04
Cost per questionAt your monthly volume
10%

Every score opens to the rows behind it

A switch verdict still lists the questions the new model got wrong. Open one to see its result next to your reference, the SQL your agent returned, and the grader's reason.

This is one of the 3 questions where Sonnet 5.5 Medium returned a wrong number with no caveat. The order counts match, so the result looks fine at a glance. The revenue doesn't.

Text-to-SQL · question 212 of 480Lost to Opus 5.5 High
QuestionQ3 net revenue and order count by region, for orders with a subscription line.
Result, your reference vs Sonnet 5.5 Medium3 of 3 rows · orders match · net_revenue differs
regionordersnet_revenuereferencenet_revenueSonnet 5.5 Mediumdiff
AMER11,2044,182,3404,960,115+18.6%
EMEA7,3882,617,9153,129,640+19.5%
APAC3,0711,094,2601,251,903+14.4%
SQL your agent returned
select o.region,       count(distinct o.order_id) as orders,       sum(o.net_amount) as net_revenuefrom analytics.orders ojoin analytics.order_items i  on i.order_id = o.order_idwhere i.product_type = 'subscription'  and not o.is_test_account  and o.order_date >= '2026-07-01'  and o.order_date < '2026-10-01'group by 1
Grader

Order counts match because the query counts distinct order IDs. Revenue runs 14 to 20% high. The query sums orders.net_amount after joining order_items, so an order with two subscription lines counts twice. The answer gave no caveat.

Result match0 of 1
Silent wrong answerYes
Assumptions statedCalendar Q3
Cost$0.031

Sonnet 5.5 ties Fable 5.1 at a fifth of the price

AnalystAgent asks the questions a data analyst answers from spreadsheets and documents. Sonnet 5.5 ties Fable 5.1 for first. Both OpenAI flagships trail by six points. Klu reruns these models on your warehouse.

AnalystAgent17 models · Oct 7, 2026
Questions right
1
Fable 5.1 Max$20 per 1M
57.5%
2
Sonnet 5.5 Max$4 per 1M
57.5%
3
Opus 5.5 Max$8 per 1M
56.2%
4
GPT-6 Astra Max$20 per 1M
51.2%
5
GPT-6.1 Sol Max$4 per 1M
50.0%
8
Kimi K3 Max$6 per 1M
38.8%

Where data teams use Klu

Each one starts with a default way to score it. Change it if yours differs.

Result set match

Text-to-SQL

Questions from Slack or your BI tool, answered with a query. Scored on whether the rows match your analyst's reference, with silent wrong answers weighted heaviest.

Row-level diff

SQL migrations

Redshift, Teradata, or stored procedures ported to Snowflake and dbt. Each translated model must return the same rows and totals as the legacy one.

Field-level match

Data cleaning

Normalize vendor names, map job titles to your taxonomy, flag duplicate accounts. Exact match per field against rows your team labeled.

Every number traced

Metric commentary

Weekly business review drafts written from your metric tables. Every figure has to trace to the source data. One invented number fails the draft.

One line per workflow, scored on the result

Klu reruns every workflow on the same frozen questions and says which to switch and which to keep. Your champion model stays until someone on your team picks a new one.

Sonnet 5.5 releaseReport ready 1 hr after release
Text-to-SQLSwitchSonnet 5.5 Medium
▲ +1.2 pts−58% cost per question
SQL migrationsKeepOpus 5.5 High
▼ 7 tables differSonnet 5.5 vs legacy output
Data cleaningKeepHaiku 4.5
→ +0.3 ptsSonnet 5.5 costs 4× more
Metric commentarySwitchSonnet 5.5 Medium
▲ +3.4 pts0 untraced numbers

Find your best model

Autopilot for peak performance, lowest price
Checked when labs ship