ProgramBench

A benchmark that asks an AI agent to rebuild a complete program, in any language, from a compiled binary and its documentation, graded by hidden behavioral tests

Agentic codingFully resolved

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 13 min read

What is ProgramBench?

ProgramBench gives a coding agent a compiled program and its documentation and asks it to write a complete codebase that behaves the same way. The agent picks the language, the architecture and the file layout. The grader never reads the code. It runs the rebuilt program against a hidden suite of behavioral tests and checks the output against the reference executable.

There are 200 tasks. The small end is command-line tools like jq and ripgrep. The large end is FFmpeg, SQLite and the PHP interpreter. Across the set, 248,853 tests do the grading. That is the count in the paper; Anthropic and the leaderboard round it to "247,000+". Meta FAIR built the benchmark with co-authors at Harvard and Stanford, including John Yang, Kilian Lieret, Ofir Press and Diyi Yang, and posted it to arXiv on 2026-05-05 as arXiv:2605.03546.

Earlier code benchmarks test a bounded edit. SWE-bench asks for a fix to one issue in a repository that already exists. Commit0, DevBench and NL2Repo-bench ask for whole repositories, but they hand the model a skeleton of classes and method signatures to fill in. ProgramBench takes away the source and the skeleton. The model has to decide how the program should be built, then grind through hundreds of edge cases until its behavior matches.

That is very hard for current models. In the May paper, nine models ran the full set and none fully resolved a single task. As of the 2026-09-30 Vals.ai update, Claude Opus 5.5 leads with 37 of 200 tasks fully resolved, or 18.5%. Most models resolve fewer than 4%.

How a task works

Each task starts in a fresh git repository with no history. The agent gets the reference executable with execute-only permission, so it can run the program and watch what it does but cannot load the binary into a disassembler like Ghidra. It also gets the program's documentation. Decompilation is off the table by design. The agent rebuilds the specification from observed behavior and the docs.

The environment is an ubuntu:22.04 Docker image with Rust 1.92.0, Python 3.12, Go 1.21.0, build-essential, cmake, git and tmux, running on 20 CPUs and 60 GB of RAM. There is no internet. Building and inference run in separate containers.

The scaffold is mini-SWE-agent, a deliberately minimal loop with a bash tool and nothing else, adopted from SWE-bench. With no custom tools in the loop, the score says more about the model than about the tooling around it. Runs stop at 1,000 steps or six hours of wall clock. Vals adds a 180-second timeout on each action.

Almost every run ends because the agent decides it is done. 98.1% of runs end by voluntary submission, 1.9% hit the time limit, and 0.05% exhaust their steps. The budget is not what holds scores down. Agents submit programs they believe are finished, and the tests disagree.

The reference programs are mostly systems code: 107 in Rust, 46 in Go, 45 in C or C++, 1 in Java and 1 in Haskell. The Hugging Face dataset card counts 33 as C. The median program has 8,635 lines in 50 files with 10 runtime dependencies, and sizes range from 212 lines to 2,701,283.

The model's choice of language shapes its score. Left alone, runs chose Python 36% of the time, Rust 25%, Go 20%, C or C++ 13% and shell 6%, and matched the reference language in exactly half of runs. When the authors forced a language different from the reference, GPT 5.4, GPT 5.4 mini and GPT 5 mini improved by 4.2%, while Claude Opus 4.7 and Opus 4.6 got worse. Free language choice is part of what ProgramBench measures, and the two vendors' models respond to it in opposite directions.

How scoring works

The primary metric is % Resolved, the share of the 200 tasks where the rebuilt program passes every hidden test and the run carries no cheating flag. Fail one test out of 770 and the task counts as unresolved. That strictness is why headline numbers sit in single digits.

Two secondary metrics show partial progress. % Almost Resolved counts tasks where at least 95% of tests pass. Raw pass rate is the mean share of hidden tests passed per task. Frontier raw pass rates on Vals sit between 82% and 87%, which sounds close to done until you set it next to fully resolved rates of 2% to 18.5% for the same models. The last few percent of behavior, the odd flag combinations and exact error strings, separate a working clone from a near miss.

The tests come from two places. Agent-driven fuzzing against the reference binary generated 79.5% of them, and the other 20.5% come from the projects' existing test suites. Six generation prompts target argument parsing, configuration, help output, I/O, subcommand dispatch and TUI interaction. Each task has a median of 770 tests, with a range of 224 to 14,645.

The generated suites reach a median 86.2% line coverage, against 64.3% for the developer-written suites. The authors also ran an assertion linter to catch weak tests, and it cut the "dummy" pass rate from 18.5% to 3.7%.

The code is MIT-licensed on GitHub, the tests are on Hugging Face, and submissions go through the ProgramBench submissions repo.

Versions and subsets

ProgramBench has no published version numbers. Every leaderboard uses the same 200 instances. What differs is who runs them and how.

The paper, from May 2026, evaluated nine models, and all nine scored 0.0% resolved. The official leaderboard, run by Meta FAIR and updated 2026-09-28, lists 25 entries and the first non-zero resolved scores. Vals.ai runs the benchmark independently and has the newest models.

Anthropic reports a third variant, a "golden" subset. It drops the 34 tasks whose reference binary scores below 0.9 on its own hidden suite, a sign of flaky tests, which leaves 166. It scores each task only against tests the reference passes and removes the six-hour limit. That is fairer to the model, since no one is penalized for a test the original program fails, but it produces a number that does not line up with either public leaderboard. The method is described in the Claude Opus 5.5 System Card, section 8.10.1, and the Claude Fable 5.1 System Card, section 8.11.1.

Current leaderboard

Four result sets exist. Each uses a different run set, metric or task subset, so each gets its own table.

Vals.ai independent run

Vals runs every model through mini-SWE-agent on its Valkyrie infrastructure with bash only, no internet, 1,000 steps, a six-hour limit and a 180-second action timeout. Score is fully resolved out of the 200 public tasks. Vals lists 62 models, most at 0.00%, with MiniMax-M2.7 in last place at rank 62. The table shows the top of the list.

Comparable within this table, since every row shares one harness; not comparable with programbench.com, which uses different runs and effort settings. Cost is the page's "Cost / Test" column, read as USD per task because its magnitudes match programbench.com's per-task costs. Vals publishes run settings only for Opus 5.5.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic18.50%mini-SWE-agent via Valkyrie; max effort, temperature 1.0, 128,000 max output tokens2026-09-30Vals. 130 almost resolved, 87.0% raw pass, $68.05/task, 2h24m, list $4/$20 per M tokens. The Vals model page lists $32.14, 79 min and 18.50% ±2.75 for the same score
Claude Fable 5.1Anthropic7.00%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals. 107 almost resolved, 82.7% raw pass, $58.52/task, 2h27m, list $10/$50
Claude Sonnet 5.5Anthropic6.50%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals. $35.86/task, 2h07m, list $2/$10
GPT-6 AstraOpenAI5.50%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals. 100 almost resolved, 85.4% raw pass, $11.57/task, 27m02s, list $10/$50
GPT-6.1 SolOpenAI3.00%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals. 92 almost resolved, 85.3% raw pass, $2.87/task, 1h10m, list $2/$10
Claude Opus 5Anthropic3.00%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals. 82.27% raw pass, $60.29/task, 2h55m, list $5/$25
Gemini 4 ArgonGoogle2.50%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals. 82 almost resolved, 83.7% raw pass, $17.66/task, 1h11m, list $4/$20
Muse Spark 1.3 MaxMeta2.50%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals. $54.29/task, 5h28m, list $1.25/$4.25
Claude Fable 5Anthropic2.00%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals. 66 almost resolved, $75.68/task, 2h37m, list $10/$50
GPT-6 SolOpenAI2.00%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals. $11.67/task, 46m11s, list $2/$10
Kimi K3Moonshot2.00%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals. No cost, duration or price shown
GPT-5.6 SolOpenAI1.50%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals. No cost, duration or price shown
Claude Opus 4.8Anthropic1.00%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals, rank 15. No cost, duration or price shown
GPT-5.6 TerraOpenAI0.50%mini-SWE-agent via Valkyrie; settings not published2026-09-30Vals, rank 20. No cost, duration or price shown
Vals.aimini-SWE-agent200 tasksFully resolvedvals.ai
ProgramBench: fully resolved vs cost per taskVals.ai, 200 tasks, mini-SWE-agent, updated 2026-09-30
10%20%$2$5$10$20$50$100
FrontierOpenAIAnthropicBehind the frontier
Cost per task, log scale, lower to the right
Source: Vals.ai ProgramBench leaderboard, updated 2026-09-30, accessed 2026-10-06. Score is the share of 200 tasks fully resolved. Cost is the leaderboard's Cost / Test column, read as USD per task with no further derivation. Vals' Opus 5.5 model page lists $32.14 and 79 minutes for the same score; the chart uses the leaderboard table.

Opus 5.5 leads Fable 5.1 by 11.5 points, and every other model on the board sits at 7.0% or below. It also gets near-misses right more often than anyone, with 130 of 200 tasks almost resolved against 107 for Fable 5.1 and 100 for GPT-6 Astra.

The cost frontier runs through five models: GPT-6.1 Sol at $2.87, GPT-6 Astra at $11.57, Claude Sonnet 5.5 at $35.86, Claude Fable 5.1 at $58.52 and Claude Opus 5.5 at $68.05. Everything else is dominated. GPT-6 Sol costs $0.10 more than Astra and scores 3.5 points lower. Gemini 4 Argon costs more than Astra and scores lower. Muse Spark 1.3 Max, Claude Opus 5 and Claude Fable 5 all cost over $54 per task and resolve 3.0% or less. The cheapest big jump on the board is the last one, where Opus 5.5 adds 11.5 points over Fable 5.1 for $9.53 more per task.

Generational gains inside one harness are large. Opus 5.5 resolves 18.5% where Opus 5 resolves 3.0%. Fable 5.1 resolves 7.0% where Fable 5 resolves 2.0%, and it costs $17.16 less per task.

Official programbench.com leaderboard

Meta FAIR runs these entries with mini-SWE-agent and no internet, and shows the effort setting in the model name. Resolved and nearly resolved are both shares of 200 tasks.

Comparable within this table; not comparable with Vals, which ran different runs and settings. Opus 5 scores 4.5% here and 3.0% on Vals, and GPT-5.6 Sol at xhigh scores 1.0% here against 1.5% on Vals. This board has no Opus 5.5, Fable 5.1, GPT-6 Astra or Gemini 4 entries.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5 (xhigh)Anthropic4.5%mini-SWE-agent, no internet, xhigh effort2026-09-28programbench.com. 37.0% nearly resolved, $50.53/task
Muse Spark 1.3 (max)Meta2.5%mini-SWE-agent, no internet, max effort2026-09-28programbench.com. 25.0% nearly resolved, $6.46/task
Muse Spark 1.3 (xhigh)Meta1.0%mini-SWE-agent, no internet, xhigh effort2026-09-28programbench.com. 16.5% nearly resolved, $2.02/task
GPT-5.6 Sol (xhigh)OpenAI1.0%mini-SWE-agent, no internet, xhigh effort2026-09-28programbench.com. 15.5% nearly resolved, $6.08/task
GPT 5.5 (xhigh)OpenAI0.5%mini-SWE-agent, no internet, xhigh effort2026-09-28programbench.com. 13.5% nearly resolved, $8.85/task
GPT 5.5 (high)OpenAI0.5%mini-SWE-agent, no internet, high effort2026-09-28programbench.com. 5.0% nearly resolved, $3.65/task
Gemini 3.6 FlashGoogle0.5%mini-SWE-agent, no internet, effort not shown2026-09-28programbench.com. 4.0% nearly resolved, $4.83/task
GPT-5.6 SolOpenAI0.5%mini-SWE-agent, no internet, effort not shown2026-09-28programbench.com. 2.5% nearly resolved, $1.00/task
Claude Opus 4.8 (xhigh)Anthropic0%mini-SWE-agent, no internet, xhigh effort2026-09-28programbench.com. 16.5% nearly resolved, $21.02/task
GLM-5.2Z.ai0%mini-SWE-agent, no internet, effort not shown2026-09-28programbench.com. 8.5% nearly resolved, $25.36/task
Gemini 3.7 FlashGoogle0%mini-SWE-agent, no internet, effort not shown2026-09-28programbench.com. 5.5% nearly resolved, $2.04/task
Muse Spark 1.2 (xhigh)Meta0%mini-SWE-agent, no internet, xhigh effort2026-09-28programbench.com. 4.5% nearly resolved, $1.48/task
Claude Opus 4.7 (xhigh)Anthropic0%mini-SWE-agent, no internet, xhigh effort2026-09-28programbench.com. 4.5% nearly resolved, $10.96/task
Muse Spark 1.1 (xhigh)Meta0%mini-SWE-agent, no internet, xhigh effort2026-09-28programbench.com. 4.0% nearly resolved, $0.73/task
Gemini 3.5 FlashGoogle0%mini-SWE-agent, no internet, effort not shown2026-09-28programbench.com. 3.0% nearly resolved, $5.60/task
Claude Opus 4.7Anthropic0%mini-SWE-agent, no internet, effort not shown2026-09-28programbench.com. 3.0% nearly resolved, $3.81/task
Claude Opus 4.6Anthropic0%mini-SWE-agent, no internet, effort not shown2026-09-28programbench.com. 2.5% nearly resolved, $11.38/task
GPT 5.5OpenAI0%mini-SWE-agent, no internet, effort not shown2026-09-28programbench.com. 1.5% nearly resolved, $1.21/task
Claude Sonnet 4.6Anthropic0%mini-SWE-agent, no internet, effort not shown2026-09-28programbench.com. 1.0% nearly resolved, $26.73/task
GPT 5.4, Gemini 3.1 Pro, Gemini 3 Flash, Claude Haiku 4.5, GPT 5.4 mini, GPT 5 miniOpenAI, Google, Anthropic0%mini-SWE-agent, no internet, effort not shown2026-09-28programbench.com, ranks 20 to 25. 0.0% nearly resolved, $0.03 to $1.51/task
Meta FAIRmini-SWE-agent200 tasksResolvedprogrambench.com

Opus 5 at xhigh leads with 4.5% resolved and 37.0% nearly resolved. The next model, Muse Spark 1.3 at max effort, nearly resolves 25.0%. Muse Spark 1.3 gets there for $6.46 per task against $50.53 for Opus 5, so it is the better buy on this board if near-misses are useful to you.

This board shows what effort settings do, which the Vals table cannot. GPT 5.5 nearly resolves 1.5% at its default, 5.0% at high and 13.5% at xhigh. Muse Spark 1.3 resolves 1.0% at xhigh and 2.5% at max. Claude Opus 4.7 nearly resolves 3.0% at default and 4.5% at xhigh. More test-time compute buys real gains on this benchmark, and a model's name alone does not tell you its score.

Anthropic golden subset

Anthropic reports a mean hidden-test pass rate on its 166-task golden subset, with mini-SWE-agent and no six-hour limit.

Comparable within this table only; vendor-reported, 166 tasks instead of 200, mean pass rate instead of fully resolved, and no time limit.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic91.2%mini-SWE-agent, 166 golden tasks, no six-hour limit2026-09-22Opus 5.5 System Card, section 8.10.1
Claude Fable 5.1Anthropic87.6%mini-SWE-agent, 166 golden tasks, no six-hour limit2026-09-01Fable 5.1 System Card, section 8.11.1; repeated in the Opus 5.5 card
Claude Fable 5Anthropic86.3%mini-SWE-agent, 166 golden tasks, no six-hour limit2026-09-01Fable 5.1 System Card, section 8.11.1
Claude Opus 5Anthropic85.4%mini-SWE-agent, 166 golden tasks, no six-hour limit2026-09-01 and 2026-09-22Both system cards
Anthropicmini-SWE-agent166 golden tasksMean hidden-test pass rate

Opus 5.5 leads at 91.2% against 87.6% for Fable 5.1. The Vals raw pass rate column above, where Opus 5.5 scores 87.0%, uses all 200 tasks and keeps the time limit, so the two pass rates measure different things.

Of the vendors checked, only Anthropic reports ProgramBench. OpenAI's GPT-6 Astra system card and the GPT-6.1 Sol addendum do not mention it. Google has not published a figure, and its Gemini 4 evaluation methodology page returned a 404 when checked on 2026-10-06. The Astra and Gemini 4 numbers on this page come only from Vals. The Opus 5.5 release post does not mention ProgramBench either; the results live in the system card.

Multi-agent harness results

Anthropic's system cards also test multi-agent systems on the 166 golden tasks, with a 1M-token limit per agent and a bash tool. They publish curves, not numbers, so the scores below are read off a chart.

Not comparable with any other table. Latency is derived from fixed prefill and decode rates plus tool time, cost assumes perfect cache hits, and the Fable 5.1 run used an internal endpoint with Opus 5 as a refusal fallback, where 72% of episodes had at least one fallback turn and under 1% of turns fell back. Anthropic says to read these as a harness comparison, not an absolute score.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Fable 5.1Anthropic0.6, read from a curveFive-agent team, 166 golden tasks, 1M tokens per agent, bash2026-09-01Fable 5.1 System Card, section 8.13.1. 2x lower latency than one agent; async subagents reach the highest final score; both multi-agent setups use more tokens
Claude Opus 5.5 (alternate snapshot)Anthropic0.6, read from a curveFive-agent team, 166 golden tasks, 1M tokens per agent, bash2026-09-22Opus 5.5 System Card, section 8.12.1. 2.7x lower latency than one agent; async subagents fall between single agent and team
AnthropicFive-agent team166 golden tasks

The result is about speed. A five-agent team reaches 0.6 between 2x and 2.7x faster than a single agent and spends more tokens doing it. If you are planning multi-agent orchestration for long builds, this is the only published ProgramBench evidence on how team structure changes wall-clock time.

Historical progression

The paper's nine models resolved nothing. The first resolved tasks appear on the September leaderboards.

The paper and programbench.com rows come from the maintainers' mini-SWE-agent runs, with vendor-default settings in the paper and effort settings on the live board. The Vals row is an independent run, so the move from 4.5% to 18.5% crosses leaderboards and is not a measured gain.

DateModelOrganizationResolvedAlmost resolvedHarness and setupSource and notes
2026-05-05Claude Opus 4.7Anthropic0.0%3.0%mini-SWE-agent, vendor-default settingsarXiv paper. $3.81/task
2026-05-05Claude Opus 4.6Anthropic0.0%2.5%mini-SWE-agent, vendor-default settingsarXiv paper. $11.38/task
2026-05-05Claude Sonnet 4.6Anthropic0.0%1.6%mini-SWE-agent, vendor-default settingsarXiv paper. $27.09/task in the paper, $26.73 on the live board
2026-05-05Gemini 3.1 Pro, GPT 5.4, GPT 5.4 mini, Claude Haiku 4.5Google, OpenAI, Anthropic0.0%0.0%mini-SWE-agent, vendor-default settingsarXiv paper. The paper's other models also scored 0.0%
2026-09-28Claude Opus 5 (xhigh)Anthropic4.5%37.0%mini-SWE-agent, no internet, xhigh effortprogrambench.com. First resolved tasks
2026-09-28Muse Spark 1.3 (max)Meta2.5%25.0%mini-SWE-agent, no internet, max effortprogrambench.com
2026-09-28GPT-5.6 Sol (xhigh)OpenAI1.0%15.5%mini-SWE-agent, no internet, xhigh effortprogrambench.com
2026-09-30Claude Opus 5.5Anthropic18.5%130 of 200 tasksmini-SWE-agent via Valkyrie, max effortVals. Model released 2026-09-22; 37 tasks resolved
mini-SWE-agentResolved

The measured gains are the ones inside a single harness. On Vals, Opus 5 resolves 3.0% and Opus 5.5 resolves 18.5%. On Anthropic's golden subset, pass rate rises from 86.3% for Fable 5 to 87.6% for Fable 5.1 to 91.2% for Opus 5.5.

Documented failure modes

Monolithic programs. Models write programs that are much smaller and flatter than the originals. Among solutions passing 75% of tests, 207 runs across 88 tasks, the median solution has 1,173 lines in 3 files. The median reference has 3,068 lines in 15 files, so solutions carry 0.38x the lines and 0.20x the files. 67% have shallower directory trees than the reference, and only 15% exceed its line count. Opus 4.7's median solution has 39 functions where the reference has 133. A program that passes the tests as one giant file is a working clone and a maintenance problem.

Opposite working styles. GPT 5.4 writes 96% of its code in a single median turn and runs a median of 17 commands. Sonnet 4.6 runs a median of 868 commands, up to 1,978 turns, at $27.09 per task. Opus 4.7 puts 67% of its code into its largest single edit. In the paper, both styles ended at 0.0% resolved.

Cheating when the internet is open. With web access, Sonnet 4.6 cheated in 36% of runs, Opus 4.6 in 21%, Gemini 3 Flash in 20% and GPT 5 mini in 1%. Looking up the original source code accounted for 79% to 95% of violations. The authors tried catching this with an LLM judge, and agreement was only moderate, with a pooled Fleiss kappa of 0.57 and a range of 0.16 to 0.60. So they blocked the internet entirely. Any internal version of this eval needs the same network isolation.

Behavioral tests are a lower bound. A passing test can still hide wrong behavior, while a failing test is definitive. A finite suite covers less than the full specification, and speed, memory and disk use are not scored at all. A resolved task means the clone matched every probe, not that it would survive production traffic.

Contamination is blocked, not measured. Every reference program is an open-source project. The paper makes no claim about memorization. Its defense is cutting off code lookup: execute-only binaries, no internet, fresh git history, and separate build and inference containers. Its test audit found zero tests that invoke undocumented flags, and 5 of 200 instances with what the paper calls "plausibly implementation-dependent output".

Documentation dependence. A task is only feasible when the executable's interface is well documented. Programs with thin docs leave the agent guessing at behavior the tests expect.

Flaky tests. On 34 tasks the reference binary scores below 0.9 on its own hidden suite. Anthropic drops those tasks; the public leaderboards keep them.

Statistical noise. Vals reports Opus 5.5 at 18.50% ±2.75. Settings for the other Vals models are unpublished, and that ±2.75 on the leader is wider than most gaps among the models between 2% and 3%.

SWE-bench fixes one issue inside an existing repository, with the codebase as a guide. ProgramBench builds the entire program with no source at all. Anthropic cites ProgramBench as a long-context test because episodes fill up to the 1M-token context window.

Commit0, DevBench and NL2Repo-bench are also from-scratch benchmarks, but they supply a skeleton with predefined classes and signatures and score against the source structure. ProgramBench grades only behavior against the reference executable, which leaves the architecture entirely to the model.

HumanEval and BigCodeBench test function-level completion from a prompt. There is no multi-file architecture and no executable to check against.

Terminal-Bench 4.0 covers shell tasks and FrontierCode covers repository-level bug and feature work. Both appear on the Opus 5.5 release post and ProgramBench does not. Their scores use different tasks and harnesses and do not compare with ProgramBench numbers. For other repository-level agent evals, see DeepSWE, FrontierSWE v2 and CursorBench. The LLM evaluation overview covers how these fit together.

What ProgramBench means for teams choosing a model

If you need whole programs built, Opus 5.5 is the only real option. It fully resolves 37 of 200 tasks on Vals, against 14 for Fable 5.1 and 11 for GPT-6 Astra, and almost resolves 130. It costs $68.05 per task and takes 2h24m.

If cost is the constraint, follow the frontier. GPT-6.1 Sol reaches 3.0% for $2.87. GPT-6 Astra resolves 5.5% for $11.57 in 27 minutes. Sonnet 5.5 resolves 6.5% for $35.86, and Fable 5.1 resolves 7.0% for $58.52. Skip Opus 5. At $60.29 it resolves 3.0%, the same as Sol at 21 times the price, and Fable 5.1 beats it by 4.0 points for $1.77 less.

Don't treat any model as production-grade for full rebuilds. Even Opus 5.5 fails 81.5% of tasks outright, and raw pass rates of 82% to 87% hide that. The almost-resolved count is a better guide to who gets close: 130 for Opus 5.5, 107 for Fable 5.1, 100 for Astra, 92 for Sol and 82 for Gemini 4 Argon. For most teams the realistic use is a near-complete draft that an engineer finishes.

Upgrade within a family. Opus 5.5 resolves 18.5% where Opus 5 resolves 3.0%. Fable 5.1 resolves 7.0% where Fable 5 resolves 2.0%, and costs $17.16 less per task. Per-token price misleads here. Fable 5.1 scores 11.5 points below Opus 5.5 at 2.5x the list price per token, $10/$50 against $4/$20, yet still costs less per task at $58.52 against $68.05.

If wall-clock time matters more than dollars, look at Astra. On Vals it finishes in 27 minutes against 2h07m to 2h27m for Sonnet 5.5, Opus 5.5 and Fable 5.1. On Anthropic's vendor-reported multi-agent harness, a five-agent team reaches 0.6, read from a curve, 2x to 2.7x faster than one agent, at higher token cost.

Pick from effort and harness, not from the model name. programbench.com and Vals disagree on the same model, with Opus 5 at 4.5% on one and 3.0% on the other, and effort settings move scores by several points. Rerun the eval at the effort level you plan to ship, and pin the date of any frontier model number you rely on.

You can compare these models across other benchmarks on the Klu LLM leaderboard. To run the same kind of behavior-graded eval on your own codebase and tasks, see how engineering teams set it up in Klu on the engineering page.