Switch models when the numbers match, not when the SQL looks right
Klu reruns your text-to-SQL and transformation evals the day a new model ships and checks every result against your analysts' reference queries.
The rows have to match. The SQL can differ.
Two correct queries rarely look alike, so Klu scores what they return. Your rubric sets what counts as the same result, like row order, rounding, and column names. A confident wrong total costs the most.
Every score opens to the rows behind it
A switch verdict still lists the questions the new model got wrong. Open one to see its result next to your reference, the SQL your agent returned, and the grader's reason.
This is one of the 3 questions where Sonnet 5.5 Medium returned a wrong number with no caveat. The order counts match, so the result looks fine at a glance. The revenue doesn't.
| region | orders | net_revenuereference | net_revenueSonnet 5.5 Medium | diff |
|---|---|---|---|---|
| AMER | 11,204 | 4,182,340 | 4,960,115 | +18.6% |
| EMEA | 7,388 | 2,617,915 | 3,129,640 | +19.5% |
| APAC | 3,071 | 1,094,260 | 1,251,903 | +14.4% |
select o.region, count(distinct o.order_id) as orders, sum(o.net_amount) as net_revenuefrom analytics.orders ojoin analytics.order_items i on i.order_id = o.order_idwhere i.product_type = 'subscription' and not o.is_test_account and o.order_date >= '2026-07-01' and o.order_date < '2026-10-01'group by 1
Order counts match because the query counts distinct order IDs. Revenue runs 14 to 20% high. The query sums orders.net_amount after joining order_items, so an order with two subscription lines counts twice. The answer gave no caveat.
Sonnet 5.5 ties Fable 5.1 at a fifth of the price
AnalystAgent asks the questions a data analyst answers from spreadsheets and documents. Sonnet 5.5 ties Fable 5.1 for first. Both OpenAI flagships trail by six points. Klu reruns these models on your warehouse.
Where data teams use Klu
Each one starts with a default way to score it. Change it if yours differs.
Text-to-SQL
Questions from Slack or your BI tool, answered with a query. Scored on whether the rows match your analyst's reference, with silent wrong answers weighted heaviest.
SQL migrations
Redshift, Teradata, or stored procedures ported to Snowflake and dbt. Each translated model must return the same rows and totals as the legacy one.
Data cleaning
Normalize vendor names, map job titles to your taxonomy, flag duplicate accounts. Exact match per field against rows your team labeled.
Metric commentary
Weekly business review drafts written from your metric tables. Every figure has to trace to the source data. One invented number fails the draft.
One line per workflow, scored on the result
Klu reruns every workflow on the same frozen questions and says which to switch and which to keep. Your champion model stays until someone on your team picks a new one.
Find your best model
Autopilot for peak performance, lowest price
Checked when labs ship