Better output for less
I build your evals, test every frontier model, and move each use case to its best model
Most teams still run last year's model
Newer models beat it on quality and cost, and nobody has re-tested
Higher quality, lower spend, and zero setup
Higher quality output
Rubrics built with your team and prompts tuned until each use case clears the bar
Lower model spend
Each use case runs on the cheapest model that clears the bar
Everything set up for you
A Klu workspace configured for your business on day one
- 3 use cases mapped
- Judges and weights set
- Caps on
- Reruns on
Two hours of your team's time
A kickoff and a review instead of 3–6 weeks of an engineer building evals
A 20-minute call
We go through what you run and what it costs, and I'll tell you if a sprint won't pay for itself
A two-week sprint
After one kickoff I build the evals, test every model, and pick a winner for each use case
A weekly retainer
An hour with me each week and a verdict within a day of every model release
Start with a sprint and keep the expert if it pays
Both plans include Klu Team, so everything I build lives in your workspace
Model Selection Sprint$1,000one timeEvery use case tested and moved to its best model, with a written report | Advisory Retainer$1,600per month · cancel anytimeYour models kept on the frontier as new ones ship, with an hour of my time every week | |
|---|---|---|
| Use cases mapped with your team | ||
| Evals designed and built | Kept current | |
| Prompts and settings tuned | ||
| Frontier models tested | Once | Every release |
| Written recommendation | One report | Weekly note |
| Time with Stephen | Kickoff + review | 1 hour every week |
| New use cases added | — | |
| Direct Slack line | — | |
| Klu Team plan | First month | Included |
Book a sprint callOne payment with no commitment | Book a retainer callStart here if you already have evals |
- Use cases mapped with your team
- Evals designed and built
- Prompts and settings tuned
- Frontier models tested
- Once
- Written recommendation
- One report
- Time with Stephen
- Kickoff + review
- New use cases added
- —
- Direct Slack line
- —
- Klu Team plan
- First month
- Use cases mapped with your team
- Evals designed and built
- Kept current
- Prompts and settings tuned
- Frontier models tested
- Every release
- Written recommendation
- Weekly note
- Time with Stephen
- 1 hour every week
- New use cases added
- Direct Slack line
- Klu Team plan
- Included
Who this is for
A good fit if
- You have AI features in production, or about to be
- Your model bill is over a few thousand dollars a month
- Nobody on the team owns model quality full time
- You want answers, not a new eval practice to run
Probably not if
- You're still prototyping and nothing runs at volume
- You already have an eval team running this every release
- You need a custom model trained from scratch
- You want us to build your product, not set up its models
Find out in 20 minutes if a sprint pays for itself
- Walk me through what you run and what it costs
- I'll point to the use case most likely to be overpaying
- No prep, no slides, no commitment

- 20 minutes
- Zoom
- Times in your time zone
Find your best model
Autopilot for peak performance, lowest price
Checked when labs ship