Harvey's Legal Agent Benchmark (LAB)

Harvey's benchmark that gives an AI agent a partner-style instruction and a closed folder of matter documents, then grades the deliverable against expert rubrics

Knowledge workMean of the two judges' task pass rates

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 14 min read

Harvey's Legal Agent Benchmark, usually shortened to LAB, scores AI agents on the kind of assignment a partner hands a junior associate. The instruction is short, about 50 words on average. It comes with a closed folder of client-matter documents. The agent has to read the folder, work out which documents matter, and write the deliverable, which is a memo, a redline, a disclosure schedule, a deposition summary or a filing. Expert-written rubric criteria then grade the files it leaves behind.

Harvey, the legal AI company, released LAB on May 6, 2026 under an MIT license. The authors are Niko Grupen, Gabe Pereyra and Julio Pereyra (announcement). Harvey built it because the legal benchmarks before it, LegalBench, CUAD, LEXam and BigLaw Bench, test short-horizon reasoning such as reading one contract or comparing cases. None of them asks an agent to build context from loose instructions across a stack of documents and then deliver files.

The grading rule is what makes LAB hard. A task scores 1 only if every rubric criterion passes. Harvey's reasoning is that a report that finds eight of ten risks "is not 80% useful; it is materially incomplete." Under that rule, frontier models pass about 90% of individual criteria and still fail most tasks. On the Artificial Analysis run, Muse Spark 1.3 at Xhigh effort passes 95.5% of criteria and completes 30.8% of tasks. That is the best result on the board, and it still leaves about seven of every ten tasks incomplete.

How a task works

Every task has four parts. There is a partner-style instruction, a work type, a list of named deliverables, and the rubric criteria, written inline. There is no gold answer file. The matter folder is a closed universe that mixes the relevant documents with peripheral ones such as emails, firm templates and procedural files, so the agent has to sort signal from clutter before it drafts anything (architecture docs).

The harness runs each task in its own Podman container started with --network=none --cap-drop=ALL. Documents mount read-only and the output folder is writable. With no network, the agent works only from what is in the folder.

The agent gets seven tools: bash, read, write, edit, glob, grep and finish. The read tool parses .docx, .xlsx, .pptx, .pdf and plain text, so the agent can open the Word files, spreadsheets, decks and PDFs a real matter folder holds. The loop ends when the agent calls finish, replies without a tool call, or hits the --max-turns limit. Harvey has not published a default max-turns value or a timeout.

The public set holds 2,010 tasks across 27 practice areas with about 114,000 rubric criteria. Corporate M&A is the largest area at 156 tasks (eval strategies doc). Anthropic's Claude Opus 5.5 System Card, section 8.14.2, gives the same task count and reports between 23 and 194 criteria per evaluated task, with a median of 56. A typical task gives the agent 56 separate ways to fail.

How scoring works

Each criterion is pass or fail. An LLM judge at temperature 0.0 decides it, looking only at the deliverable files that criterion declares and with no golden reference to compare against. The task score is all-pass, 1 if every criterion passes and 0 otherwise. Reports also give a criterion pass rate.

Read the two numbers together. All-pass says how often the agent hands over a complete work product. Criterion pass says how close it gets. A model at 93% criterion pass and 10% all-pass is close on almost everything and finishes almost nothing, which is a different product from one that nails a few task types and fails the rest.

Because LAB has no executable ground truth, the judge is part of the measurement. Since lab-core v1.1.0 on September 18, 2026, Harvey's default averages two judges, Claude Sonnet 4.6 and GPT-5.5. The dual_all_pass_rate metric is the mean of the two judges' per-task values, so a task the judges split on scores 0.5. Artificial Analysis grades with Gemini 3.1 Pro alone. A score produced under one judge setup does not transfer to another.

Versions and who runs it

LAB has grown since launch, and one release changed how scores are computed. The table lists the changes from Harvey's GitHub releases and docs.

ReleaseDateWhat changedEffect on scores
Launch2026-05-061,250+ tasks, 24 practice areas, 75,000+ rubric criteriaBaseline set
lab-core v1.1.02026-09-18Dual-judge evaluation (Claude Sonnet 4.6 and GPT-5.5) made the default; explicit finish tool made the default end; rubric fixesTask scores can now be 0.5; earlier single-judge runs are not comparable
v1.2.02026-10-01Support for Claude 5.5 and GPT-6 modelsNew models runnable on the official harness
Current public setAs of 2026-10-012,010 tasks, 27 practice areas, about 114,000 criteriaLarger public pool
Held-out setNot published120 tasks Harvey keeps privateBasis of the Vals and Artificial Analysis leaderboards

Three outside groups publish LAB leaderboards, and each makes different choices.

Vals AI runs the held-out set on Valkyrie, its own framework for agentic benchmarks. It uses Harvey's default judge pair, GPT 5.5 at medium reasoning and Claude Sonnet 4.6, and reports the mean of the two judges' task pass rates. Its page lists 74 models. Vals has not published the task count, but every score is a multiple of 1/240, which matches 120 tasks scored by two judges.

Artificial Analysis launched Harvey LAB-AA on July 7, 2026. It runs 120 private tasks across 24 practice areas on Stirrup, its open-source harness, with one judge, Gemini 3.1 Pro. The agent gets a simple code-execution tool and nothing else, Harvey's document-generation scripts are excluded, and filenames must match exactly. AA calls this an independent reimplementation that differs from the original. It publishes all-pass, criterion pass, cost per task, time and turns, and puts the reasoning effort in each model's name.

Mercor, which its page names as Harvey's partner on LAB, publishes a leaderboard on 1,749 tasks with all-pass pass@1 and confidence intervals. It names neither the judge model nor the harness version, shows no date and has no cost data.

Current leaderboard

Each of the four result sets below uses a different task set, harness or judge, so each gets its own table.

Vals AI, held-out set, dual judge

Vals runs every model through Valkyrie on the held-out set and grades with GPT 5.5 and Claude Sonnet 4.6. It does not publish reasoning effort, turn limits or timeouts per model. Cost is the page's "Cost / Test" column, the total cost of one task run.

Comparable within this table only; not comparable with Artificial Analysis, which uses a different harness, tool set and judge.

ModelOrganizationScoreHarness and setupDateSource and notes
Muse Spark 1.2Meta25.42%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 1. $2.09/task, 24m16s, list $1.25/$4.25 per M tokens
Muse Spark 1.3 MaxMeta23.75%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 2. $2.26/task, 14m43s, list $1.25/$4.25
Muse Spark 1.3Meta22.92%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 3. $2.76/task, 17m04s, list $1.25/$4.25
Muse Spark 1.1Meta20.00%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 4. $0.80/task, 12m21s, list $1.25/$4.25
Gemini 4 ArgonGoogle19.58%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 5. $10.55/task, 43m55s, list $4/$20. Google's launch evaluation cites this run
Grok 4.6xAI15.83%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 6. $4.01/task, 45m05s, list $2/$6
Grok 4.5xAI12.92%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 7. $2.02/task, 10m15s, list $2/$6. Tied with Kimi K3
Kimi K3Moonshot AI12.92%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 8. $3.83/task, 13m58s, list $3/$15
Grok 4.7xAI12.50%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 9. $11.13/task, 42m05s, list $2/$6
MiMo V2.6 FlashXiaomi11.25%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 10. $0.09/task, 17m12s, list $0.14/$0.28
Claude Fable 5Anthropic11.25%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 12
Claude Opus 4.8Anthropic9.58%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank 16
Claude Opus 5Anthropic6.67%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank not in extract. $23.67/task, 55m38s
Claude Fable 5.1Anthropic6.67%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank not in extract. $46.21/task, 1h52m
GPT-6 AstraOpenAI5.42%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank not in extract. $26.16/task, 24m49s
Claude Opus 5.5Anthropic3.75%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank not in extract. $21.38/task, 52m55s
GPT-5.6 SolOpenAI2.50%Valkyrie, held-out set, dual judge; effort not published2026-10-02Vals, rank not in extract. $10.37/task, 22m50s
Vals AIValkyrieHeld-out set, dual judgeMean of the two judges' task pass ratesvals.ai
Harvey LAB: task pass rate vs cost per taskVals AI, held-out set, Valkyrie harness, dual judge, updated 2026-10-02
10%20%$0.1$0.2$0.5$1$2$5$10$20
FrontierOtherMetaBehind the frontier
Cost per task, log scale, lower to the right
Source: Vals AI Harvey LAB leaderboard, updated 2026-10-02, accessed 2026-10-06. Score is the task pass rate, the mean of the GPT 5.5 and Claude Sonnet 4.6 judges' all-pass rates on the held-out set. Cost is the leaderboard's Cost / Test column in USD, used as published. Claude Fable 5.1 ($46.21, 6.67%) is omitted to keep the chart at ten points. Vals does not publish reasoning effort per model.

Meta holds the top four places. Muse Spark 1.2 leads at 25.42%, 1.67 points ahead of Muse Spark 1.3 Max, and costs $2.09 per task. Gemini 4 Argon is the best model from another lab at 19.58%, 5.84 points behind, and it costs $10.55 per task, five times as much. Google's Gemini 4 Argon evaluation sources its LAB result from this Vals run rather than from a Google harness.

The cost frontier runs through three models: MiMo V2.6 Flash at $0.09 and 11.25%, Muse Spark 1.1 at $0.80 and 20.00%, and Muse Spark 1.2 at $2.09 and 25.42%. Every other model costs more for a lower score. Anthropic and OpenAI flagships sit at the bottom of the chart, at the expensive end. Claude Opus 5, Claude Opus 5.5 and GPT-6 Astra each cost over $21 per task and complete under 7% of tasks. Opus 5.5 scores below Opus 5 here, 3.75% against 6.67%.

At this resolution small gaps are noise. One task is worth 0.42 points on the dual-judge mean, which is why Grok 4.5 and Kimi K3 tie exactly at 12.92%.

Artificial Analysis, Stirrup harness, single judge

AA runs 120 private tasks on its own harness with a code-execution tool, Gemini 3.1 Pro as the only judge, and strict filename matching. It lists 55 models. The table shows the 14 the page selects by default, with values read from the page's embedded data. The page shows no snapshot date, so the date column gives the fetch date and the notes give each model's release date.

Comparable within this table only; not comparable with Vals, which uses Harvey's dual judge and a different harness.

ModelOrganizationScoreHarness and setupDateSource and notes
Muse Spark 1.3 (Xhigh)Meta30.8%Stirrup, 120 private tasks, Gemini 3.1 Pro judge, XhighFetched 2026-10-06AA. 95.5% criterion pass, 84.2 turns/task. Released 2026-09-02
Kimi K3 (Max)Moonshot AI26.7%Stirrup, 120 private tasks, Gemini 3.1 Pro judge, MaxFetched 2026-10-06AA. 94.6% criterion pass, 94.2 turns/task. Released 2026-07-16
Claude Fable 5.1 (Xhigh)Anthropic16.7%Stirrup, 120 private tasks, Gemini 3.1 Pro judge, XhighFetched 2026-10-06AA. 93.3% criterion pass, 75.3 turns/task. Released 2026-09-01
Step 5 PreviewStepFun15.8%Stirrup, 120 private tasks, Gemini 3.1 Pro judgeFetched 2026-10-06AA. 93.4% criterion pass, 95.9 turns/task. Released 2026-09-18
Claude Fable 5.1 (Max)Anthropic14.2%Stirrup, 120 private tasks, Gemini 3.1 Pro judge, MaxFetched 2026-10-06AA. 93.0% criterion pass, 82.8 turns/task. Released 2026-09-01
Claude Fable 5 (Max, Opus 4.8 fallback)Anthropic14.2%Stirrup, 120 private tasks, Gemini 3.1 Pro judge, MaxFetched 2026-10-06AA. 93.6% criterion pass, 64.1 turns/task. Released 2026-06-09
Claude Fable 5.1 (High)Anthropic13.3%Stirrup, 120 private tasks, Gemini 3.1 Pro judge, HighFetched 2026-10-06AA. 93.0% criterion pass, 63.5 turns/task. Released 2026-09-01
Claude Opus 5 (Max)Anthropic10.8%Stirrup, 120 private tasks, Gemini 3.1 Pro judge, MaxFetched 2026-10-06AA. 93.5% criterion pass, 77.3 turns/task. Released 2026-07-24
Claude Sonnet 5.5 (Max)Anthropic10.0%Stirrup, 120 private tasks, Gemini 3.1 Pro judge, MaxFetched 2026-10-06AA. 93.1% criterion pass, 173.1 turns/task. Released 2026-09-28
Qwen3.8 Max (0902)Alibaba10.0%Stirrup, 120 private tasks, Gemini 3.1 Pro judgeFetched 2026-10-06AA. 93.6% criterion pass, 142.0 turns/task. Released 2026-09-02
Claude Opus 5.5 (Max)Anthropic8.3%Stirrup, 120 private tasks, Gemini 3.1 Pro judge, MaxFetched 2026-10-06AA. 91.2% criterion pass, 115.3 turns/task. Released 2026-09-22. Anthropic's system card, section 8.14.2, reports the same values
MiniMax-M3MiniMax6.7%Stirrup, 120 private tasks, Gemini 3.1 Pro judgeFetched 2026-10-06AA. 88.4% criterion pass, 63.5 turns/task. Released 2026-06-01
Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIA3.3%Stirrup, 120 private tasks, Gemini 3.1 Pro judgeFetched 2026-10-06AA. 81.7% criterion pass, 22.1 turns/task. Released 2026-06-04
Mistral Medium 3.5Mistral0.8%Stirrup, 120 private tasks, Gemini 3.1 Pro judgeFetched 2026-10-06AA. 69.1% criterion pass, 28.4 turns/task. Released 2026-04-29
Artificial AnalysisStirrup120 private tasks, Gemini 3.1 Pro judgeAll-passartificialanalysis.ai

Muse Spark 1.3 at Xhigh leads at 30.8%, 4.1 points ahead of Kimi K3 at Max and 14.1 points ahead of the best Anthropic entry, Claude Fable 5.1 at Xhigh. The two independent boards agree that Meta's Muse Spark line is the strongest on this benchmark. Kimi K3 is the clearest disagreement between them. It places second here and eighth on Vals.

Effort settings move scores more than you would guess. Fable 5.1 scores 16.7% at Xhigh, 14.2% at Max and 13.3% at High, so the most expensive setting is not the best one for this workload. Opus 5.5 at Max scores 8.3%, below Opus 5 at Max on 10.8%, and takes 115.3 turns per task against 77.3. That matches the Vals result. Two separate harnesses with different judges both put Opus 5.5 behind Opus 5 on legal-agent work.

Turn counts say something about working style. Sonnet 5.5 at Max averages 173.1 turns per task and Qwen3.8 Max averages 142.0. Both land at 10.0%. Muse Spark 1.3 finishes in 84.2 turns and scores three times higher.

Harvey-reported

Harvey published one LAB score for a single model on its blog.

Not comparable with any other table; Harvey does not state the harness, judges or task set.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5Anthropic11.7%Not disclosed; judges and set not given2026-07-24Harvey. All-pass only, no criterion pass. Harvey calls it "a meaningful step up relative to prior Opus models" and reports parity with Opus 4.8 at max using lower reasoning levels and 26% fewer tokens
HarveyAll-passharvey.ai

AA's own run of Opus 5 at Max is 10.8%, close to Harvey's figure, but nothing in Harvey's post ties the two runs together.

Mercor OSS leaderboard

Mercor's board uses the largest task set of the three, and it is the only one that reports confidence intervals.

Comparable within this table only; its 1,749-task set and unnamed judge differ from both independent boards, and its ranking conflicts with them.

ModelOrganizationScoreHarness and setupDateSource and notes
Opus 5 (Max)Anthropic17.8% ±1.5Open-source LAB harness, version not named; unnamed LLM judgeNot shown, fetched 2026-10-06Mercor, rank 1. All-pass pass@1. Not stated who ran the models
Fable 5.1 (Max)Anthropic17.4% ±4.2Open-source LAB harness, version not named; unnamed LLM judgeNot shown, fetched 2026-10-06Mercor, rank 2
Muse Spark 1.3 (Max)Meta15.0% ±4.1Open-source LAB harness, version not named; unnamed LLM judgeNot shown, fetched 2026-10-06Mercor, rank 3
Kimi K3 (Max)Moonshot AI14.9% ±1.4Open-source LAB harness, version not named; unnamed LLM judgeNot shown, fetched 2026-10-06Mercor, rank 4
GLM-5.3 (Max)Not given14.6% ±3.8Open-source LAB harness, version not named; unnamed LLM judgeNot shown, fetched 2026-10-06Mercor, rank 5
MercorOpen-source LAB harness1,749 tasksAll-pass pass@1mercor.com

Mercor puts Opus 5 first at 17.8%, against 6.67% on Vals and 10.8% on AA, and puts Muse Spark 1.3 third. Fable 5.1, Muse Spark 1.3 and GLM-5.3 carry intervals of ±3.8 to ±4.2 points, wider than the 2.8-point spread between them, so even on its own terms this board does not separate them. I would not pick a model from it. It names no judge and no harness version, it shows no date, and it does not say who ran the models.

Historical results

Each run below used its own harness or judging, so scores do not compare across rows. They are useful for what each run found, not as a time series.

DateRunHarness and judgingWhat it found
2026-05-26Harvey initial resultsHarvey's own harness; holdout mirroring the public set; multi-judge average, judges not namedClaude Opus 4.7 led at 7.1%. Frontier models "complete less than 10% of tasks end-to-end in aggregate"
2026-07-07Artificial Analysis launchStirrup, single judge, strict filename matching28 models. Claude Fable 5 led at 14.2%. Cost per task ran from $0.02 to $18.90, about 950x
2026-07-24Harvey Opus 5 postNot disclosedOpus 5 at 11.7% all-pass
2026-08-20Harvey Tenet post-trainingHarvey's official hold-out set plus sub-benchmarks; LLM judges with granular and holistic rubric termsPost-trained Kimi K3 completes "almost twice as many" held-out tasks as base K3, a 9-point all-pass gain

The first published results, from Harvey's May 26 post, are the clearest baseline.

Not comparable with any current table; Harvey's own harness and unnamed judges differ from Vals, AA and Mercor.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 4.7Anthropic7.1%Harvey harness; holdout mirroring public distribution; multi-judge average2026-05-26Harvey. About $50.90/task, about 22 min
Claude Sonnet 4.6Anthropic5.4%Same2026-05-26Harvey
Claude Opus 4.6Anthropic4.2%Same2026-05-26Harvey
GPT-5.5OpenAI2.1%Same2026-05-26Harvey. About $17/task
Gemini 3.5 FlashGoogle0.8%Same2026-05-26Harvey. Under 6 min/task
HarveyHarvey harnessHoldout mirroring public distributionharvey.ai

The one real time series is on the AA harness, which kept the same setup from launch to now. The top score there went from 14.2% for Claude Fable 5 in July to 30.8% for Muse Spark 1.3 today, more than double in three months. Fable 5's 14.2% is the same figure the current board shows for Claude Fable 5 (Max, Opus 4.8 fallback).

The Tenet post tested Harvey's own post-trained models, and each gain is against that model's base. Beyond the Kimi K3 result, the LAB Contracts 50-task hold-out improved 20% over base, a 2-point all-pass gain. On LAB Diligence, run in an RLM harness, post-trained GLM-5.2 reached 60.1% criterion pass against a 46.1% baseline. LAB Firm Knowledge gained 15% in criterion pass with a 90% cost reduction. Harvey says it used "no customer data" for this training. The takeaway is that domain post-training moves LAB scores a lot, which matters if you are comparing a general model to a legal-specialist one.

Documented failure modes

Competence without completion. The all-pass rule exposes a gap that criterion scores hide. On AA, Muse Spark 1.3 passes 95.5% of criteria and completes 30.8% of tasks. Opus 5.5 passes 91.2% and completes 8.3%. Twelve of the 14 models AA shows by default pass between 88.4% and 95.5% of criteria, and none of them completes more than 30.8% of tasks. Vals says it directly: a model "can satisfy most individual criteria and still miss task resolution credit."

Drafting without checking. Harvey's May analysis measured which agent behaviors moved all-pass scores. Verify-and-revise loops added 1.5 points, post-draft validation 0.8, thorough research before drafting 0.4, and targeted retrieval, structured analysis in code and grounding against source documents 0.3 each. Drafting without review cost 1.2 points, and firing five or more parallel tool calls cost 0.5. The pattern is consistent. Agents that read carefully and check their own work score higher, and agents that rush to a draft score lower.

Uneven practice-area strength. No single model led every area in that May run. GPT-5.5 led the regulated and emerging-company groups, Opus 4.7 led corporate transactions and funds, and Sonnet 4.6 led privacy, tax and private client. A headline score averages over these differences.

Cost and time grow with capability. AA found a 950x spread in per-task cost at launch, and stronger models took more turns and more time. On Vals, the slowest flagship runs take close to an hour per task, and Claude Fable 5.1 takes 1h52m.

Judge dependence. LAB has no answer key, so a different judge produces a different score. Harvey's default averages Sonnet 4.6 and GPT-5.5; AA uses Gemini 3.1 Pro alone.

Vendor ownership. Harvey sells legal AI and built the benchmark. LawNext describes LAB as "built by a market participant," and commentator Houfu Ang calls vendor-led legal open source "open source theatre," pointing to minimal outside contributors.

Contamination. The 2,010 public tasks, their rubrics and the harness have been MIT-licensed on GitHub since May 6, 2026, so any model trained after that date has had the chance to see them. The Vals and AA boards score the 120-task private hold-out instead. No contamination audit has been published.

Coverage gaps. Harvey says the initial release covers representative practice areas and leaves many task families uncovered. Future versions are planned to add all BigLaw practices, in-house work and adjacent domains.

Missing settings. Vals has not published reasoning effort, turn limits or timeouts per model. Harvey has not published a default max-turns value or timeout. AA has not published its turn limit or snapshot date. Mercor has not published its judge model, harness version or date. Each gap makes a leaderboard row harder to reproduce.

GDPval-AA, also run by Artificial Analysis, takes 220 tasks from OpenAI's GDPval gold set across 44 occupations and nine industries, gives the agent a shell and web browsing, and ranks models by Elo from blind pairwise comparisons anchored to DeepSeek V4.1 Flash at max effort at 1,600. Anthropic reports Opus 5.5 at 1846 at max and 1820 at xhigh in system card section 8.14.3. GDPval-AA is broad and relative. It asks which of two outputs is better across many jobs. LAB is narrow and absolute, and asks whether one legal deliverable meets every expert criterion.

AA-Briefcase v1.1 covers multi-week projects with thousands of source files and grades with a rubric plus a pairwise panel. Anthropic reports Opus 5.5 at 1822 Elo at max in section 8.14.4. Briefcase tests sustained work across far more material; LAB tests one bounded matter in one folder, where the agent still has to choose what to pull into its context window.

SWE-bench and GAIA grade against objective checks, unit tests for SWE-bench and exact answers for GAIA. LAB has no executable ground truth and relies on LLM judges reading against expert rubrics, which is why the judge pair matters so much more here than on a coding benchmark.

The earlier legal benchmarks Harvey names, LegalBench, CUAD, LEXam and BigLaw Bench, test short-horizon reasoning without file deliverables. LAB is what you use when you want to know whether an agent can finish the job, not whether it can answer a question about a clause.

Klu's other agent benchmark pages cover Terminal-Bench 4.0, AutomationBench, Agents' Last Exam and HealthBench. For background on how evaluations like these are built, see LLM evaluation.

What it means for teams choosing a model

Start with Muse Spark if accuracy and cost both matter. It leads both independent boards, at 25.42% for Muse Spark 1.2 at $2.09 per task on Vals and 30.8% for Muse Spark 1.3 at Xhigh on AA. Mercor ranks Opus 5 and Fable 5.1 above Muse Spark 1.3, but Mercor uses a different task set and does not say who ran the models. The lead holds on the two independent runs. Gemini 4 Argon is the top lab flagship on Vals at 19.58%, at five times Muse Spark 1.2's cost for 5.8 fewer points.

Do not assume a newer flagship is better at legal work. Opus 5.5 scores below Opus 5 on both independent harnesses, 3.75% against 6.67% on Vals and 8.3% against 10.8% on AA, and takes 115 turns per task on AA against 77. If you run Opus 5 for legal agents today, test before you upgrade.

Check what the premium tier buys. Claude Fable 5.1 costs $46.21 per task on Vals for 6.67%, the same score as Opus 5 at $23.67. The extra $22.54 per task buys nothing there. On AA, Fable 5.1 does best at Xhigh, 16.7%, not at Max, 14.2%. More test-time compute is not automatically better on this workload.

For volume, look at MiMo V2.6 Flash. It scores 11.25% at $0.09 per task on Vals, within 1.25 points of Grok 4.7 at $11.13, for under 1% of the cost.

Plan for review of every output. A 90%-plus criterion pass rate does not mean a usable deliverable. On AA, the best model still leaves seven of ten tasks incomplete. In Harvey's May data the behaviors that helped most were verify-and-revise loops and post-draft validation, and drafting without review cost the most. Configure your agent to check its own work, then have an attorney check it again.

Rerun the eval on your own matters. The boards disagree on rank, judges differ, and effort settings move scores by several points. Pin the harness, judge and effort level you plan to ship, and treat any frontier model number as dated to the day it was fetched.

You can compare these models across other benchmarks on the Klu LLM leaderboard. To run the same kind of rubric-graded eval on your own contracts and matter files, see how legal teams set it up in Klu on the legal page.