What is ProgramBench?
ProgramBench gives a coding agent a compiled program and its documentation and asks it to write a complete codebase that behaves the same way. The agent picks the language, the architecture and the file layout. The grader never reads the code. It runs the rebuilt program against a hidden suite of behavioral tests and checks the output against the reference executable.
There are 200 tasks. The small end is command-line tools like jq and ripgrep. The large end is FFmpeg, SQLite and the PHP interpreter. Across the set, 248,853 tests do the grading. That is the count in the paper; Anthropic and the leaderboard round it to "247,000+". Meta FAIR built the benchmark with co-authors at Harvard and Stanford, including John Yang, Kilian Lieret, Ofir Press and Diyi Yang, and posted it to arXiv on 2026-05-05 as arXiv:2605.03546.
Earlier code benchmarks test a bounded edit. SWE-bench asks for a fix to one issue in a repository that already exists. Commit0, DevBench and NL2Repo-bench ask for whole repositories, but they hand the model a skeleton of classes and method signatures to fill in. ProgramBench takes away the source and the skeleton. The model has to decide how the program should be built, then grind through hundreds of edge cases until its behavior matches.
That is very hard for current models. In the May paper, nine models ran the full set and none fully resolved a single task. As of the 2026-09-30 Vals.ai update, Claude Opus 5.5 leads with 37 of 200 tasks fully resolved, or 18.5%. Most models resolve fewer than 4%.
How a task works
Each task starts in a fresh git repository with no history. The agent gets the reference executable with execute-only permission, so it can run the program and watch what it does but cannot load the binary into a disassembler like Ghidra. It also gets the program's documentation. Decompilation is off the table by design. The agent rebuilds the specification from observed behavior and the docs.
The environment is an ubuntu:22.04 Docker image with Rust 1.92.0, Python 3.12, Go 1.21.0, build-essential, cmake, git and tmux, running on 20 CPUs and 60 GB of RAM. There is no internet. Building and inference run in separate containers.
The scaffold is mini-SWE-agent, a deliberately minimal loop with a bash tool and nothing else, adopted from SWE-bench. With no custom tools in the loop, the score says more about the model than about the tooling around it. Runs stop at 1,000 steps or six hours of wall clock. Vals adds a 180-second timeout on each action.
Almost every run ends because the agent decides it is done. 98.1% of runs end by voluntary submission, 1.9% hit the time limit, and 0.05% exhaust their steps. The budget is not what holds scores down. Agents submit programs they believe are finished, and the tests disagree.
The reference programs are mostly systems code: 107 in Rust, 46 in Go, 45 in C or C++, 1 in Java and 1 in Haskell. The Hugging Face dataset card counts 33 as C. The median program has 8,635 lines in 50 files with 10 runtime dependencies, and sizes range from 212 lines to 2,701,283.
The model's choice of language shapes its score. Left alone, runs chose Python 36% of the time, Rust 25%, Go 20%, C or C++ 13% and shell 6%, and matched the reference language in exactly half of runs. When the authors forced a language different from the reference, GPT 5.4, GPT 5.4 mini and GPT 5 mini improved by 4.2%, while Claude Opus 4.7 and Opus 4.6 got worse. Free language choice is part of what ProgramBench measures, and the two vendors' models respond to it in opposite directions.
How scoring works
The primary metric is % Resolved, the share of the 200 tasks where the rebuilt program passes every hidden test and the run carries no cheating flag. Fail one test out of 770 and the task counts as unresolved. That strictness is why headline numbers sit in single digits.
Two secondary metrics show partial progress. % Almost Resolved counts tasks where at least 95% of tests pass. Raw pass rate is the mean share of hidden tests passed per task. Frontier raw pass rates on Vals sit between 82% and 87%, which sounds close to done until you set it next to fully resolved rates of 2% to 18.5% for the same models. The last few percent of behavior, the odd flag combinations and exact error strings, separate a working clone from a near miss.
The tests come from two places. Agent-driven fuzzing against the reference binary generated 79.5% of them, and the other 20.5% come from the projects' existing test suites. Six generation prompts target argument parsing, configuration, help output, I/O, subcommand dispatch and TUI interaction. Each task has a median of 770 tests, with a range of 224 to 14,645.
The generated suites reach a median 86.2% line coverage, against 64.3% for the developer-written suites. The authors also ran an assertion linter to catch weak tests, and it cut the "dummy" pass rate from 18.5% to 3.7%.
The code is MIT-licensed on GitHub, the tests are on Hugging Face, and submissions go through the ProgramBench submissions repo.
Versions and subsets
ProgramBench has no published version numbers. Every leaderboard uses the same 200 instances. What differs is who runs them and how.
The paper, from May 2026, evaluated nine models, and all nine scored 0.0% resolved. The official leaderboard, run by Meta FAIR and updated 2026-09-28, lists 25 entries and the first non-zero resolved scores. Vals.ai runs the benchmark independently and has the newest models.
Anthropic reports a third variant, a "golden" subset. It drops the 34 tasks whose reference binary scores below 0.9 on its own hidden suite, a sign of flaky tests, which leaves 166. It scores each task only against tests the reference passes and removes the six-hour limit. That is fairer to the model, since no one is penalized for a test the original program fails, but it produces a number that does not line up with either public leaderboard. The method is described in the Claude Opus 5.5 System Card, section 8.10.1, and the Claude Fable 5.1 System Card, section 8.11.1.
Current leaderboard
Four result sets exist. Each uses a different run set, metric or task subset, so each gets its own table.
Vals.ai independent run
Vals runs every model through mini-SWE-agent on its Valkyrie infrastructure with bash only, no internet, 1,000 steps, a six-hour limit and a 180-second action timeout. Score is fully resolved out of the 200 public tasks. Vals lists 62 models, most at 0.00%, with MiniMax-M2.7 in last place at rank 62. The table shows the top of the list.
Comparable within this table, since every row shares one harness; not comparable with programbench.com, which uses different runs and effort settings. Cost is the page's "Cost / Test" column, read as USD per task because its magnitudes match programbench.com's per-task costs. Vals publishes run settings only for Opus 5.5.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 18.50% | mini-SWE-agent via Valkyrie; max effort, temperature 1.0, 128,000 max output tokens | 2026-09-30 | Vals. 130 almost resolved, 87.0% raw pass, $68.05/task, 2h24m, list $4/$20 per M tokens. The Vals model page lists $32.14, 79 min and 18.50% ±2.75 for the same score |
| Claude Fable 5.1 | Anthropic | 7.00% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals. 107 almost resolved, 82.7% raw pass, $58.52/task, 2h27m, list $10/$50 |
| Claude Sonnet 5.5 | Anthropic | 6.50% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals. $35.86/task, 2h07m, list $2/$10 |
| GPT-6 Astra | OpenAI | 5.50% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals. 100 almost resolved, 85.4% raw pass, $11.57/task, 27m02s, list $10/$50 |
| GPT-6.1 Sol | OpenAI | 3.00% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals. 92 almost resolved, 85.3% raw pass, $2.87/task, 1h10m, list $2/$10 |
| Claude Opus 5 | Anthropic | 3.00% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals. 82.27% raw pass, $60.29/task, 2h55m, list $5/$25 |
| Gemini 4 Argon | 2.50% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals. 82 almost resolved, 83.7% raw pass, $17.66/task, 1h11m, list $4/$20 | |
| Muse Spark 1.3 Max | Meta | 2.50% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals. $54.29/task, 5h28m, list $1.25/$4.25 |
| Claude Fable 5 | Anthropic | 2.00% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals. 66 almost resolved, $75.68/task, 2h37m, list $10/$50 |
| GPT-6 Sol | OpenAI | 2.00% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals. $11.67/task, 46m11s, list $2/$10 |
| Kimi K3 | Moonshot | 2.00% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals. No cost, duration or price shown |
| GPT-5.6 Sol | OpenAI | 1.50% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals. No cost, duration or price shown |
| Claude Opus 4.8 | Anthropic | 1.00% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals, rank 15. No cost, duration or price shown |
| GPT-5.6 Terra | OpenAI | 0.50% | mini-SWE-agent via Valkyrie; settings not published | 2026-09-30 | Vals, rank 20. No cost, duration or price shown |
Opus 5.5 leads Fable 5.1 by 11.5 points, and every other model on the board sits at 7.0% or below. It also gets near-misses right more often than anyone, with 130 of 200 tasks almost resolved against 107 for Fable 5.1 and 100 for GPT-6 Astra.
The cost frontier runs through five models: GPT-6.1 Sol at $2.87, GPT-6 Astra at $11.57, Claude Sonnet 5.5 at $35.86, Claude Fable 5.1 at $58.52 and Claude Opus 5.5 at $68.05. Everything else is dominated. GPT-6 Sol costs $0.10 more than Astra and scores 3.5 points lower. Gemini 4 Argon costs more than Astra and scores lower. Muse Spark 1.3 Max, Claude Opus 5 and Claude Fable 5 all cost over $54 per task and resolve 3.0% or less. The cheapest big jump on the board is the last one, where Opus 5.5 adds 11.5 points over Fable 5.1 for $9.53 more per task.
Generational gains inside one harness are large. Opus 5.5 resolves 18.5% where Opus 5 resolves 3.0%. Fable 5.1 resolves 7.0% where Fable 5 resolves 2.0%, and it costs $17.16 less per task.
Official programbench.com leaderboard
Meta FAIR runs these entries with mini-SWE-agent and no internet, and shows the effort setting in the model name. Resolved and nearly resolved are both shares of 200 tasks.
Comparable within this table; not comparable with Vals, which ran different runs and settings. Opus 5 scores 4.5% here and 3.0% on Vals, and GPT-5.6 Sol at xhigh scores 1.0% here against 1.5% on Vals. This board has no Opus 5.5, Fable 5.1, GPT-6 Astra or Gemini 4 entries.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5 (xhigh) | Anthropic | 4.5% | mini-SWE-agent, no internet, xhigh effort | 2026-09-28 | programbench.com. 37.0% nearly resolved, $50.53/task |
| Muse Spark 1.3 (max) | Meta | 2.5% | mini-SWE-agent, no internet, max effort | 2026-09-28 | programbench.com. 25.0% nearly resolved, $6.46/task |
| Muse Spark 1.3 (xhigh) | Meta | 1.0% | mini-SWE-agent, no internet, xhigh effort | 2026-09-28 | programbench.com. 16.5% nearly resolved, $2.02/task |
| GPT-5.6 Sol (xhigh) | OpenAI | 1.0% | mini-SWE-agent, no internet, xhigh effort | 2026-09-28 | programbench.com. 15.5% nearly resolved, $6.08/task |
| GPT 5.5 (xhigh) | OpenAI | 0.5% | mini-SWE-agent, no internet, xhigh effort | 2026-09-28 | programbench.com. 13.5% nearly resolved, $8.85/task |
| GPT 5.5 (high) | OpenAI | 0.5% | mini-SWE-agent, no internet, high effort | 2026-09-28 | programbench.com. 5.0% nearly resolved, $3.65/task |
| Gemini 3.6 Flash | 0.5% | mini-SWE-agent, no internet, effort not shown | 2026-09-28 | programbench.com. 4.0% nearly resolved, $4.83/task | |
| GPT-5.6 Sol | OpenAI | 0.5% | mini-SWE-agent, no internet, effort not shown | 2026-09-28 | programbench.com. 2.5% nearly resolved, $1.00/task |
| Claude Opus 4.8 (xhigh) | Anthropic | 0% | mini-SWE-agent, no internet, xhigh effort | 2026-09-28 | programbench.com. 16.5% nearly resolved, $21.02/task |
| GLM-5.2 | Z.ai | 0% | mini-SWE-agent, no internet, effort not shown | 2026-09-28 | programbench.com. 8.5% nearly resolved, $25.36/task |
| Gemini 3.7 Flash | 0% | mini-SWE-agent, no internet, effort not shown | 2026-09-28 | programbench.com. 5.5% nearly resolved, $2.04/task | |
| Muse Spark 1.2 (xhigh) | Meta | 0% | mini-SWE-agent, no internet, xhigh effort | 2026-09-28 | programbench.com. 4.5% nearly resolved, $1.48/task |
| Claude Opus 4.7 (xhigh) | Anthropic | 0% | mini-SWE-agent, no internet, xhigh effort | 2026-09-28 | programbench.com. 4.5% nearly resolved, $10.96/task |
| Muse Spark 1.1 (xhigh) | Meta | 0% | mini-SWE-agent, no internet, xhigh effort | 2026-09-28 | programbench.com. 4.0% nearly resolved, $0.73/task |
| Gemini 3.5 Flash | 0% | mini-SWE-agent, no internet, effort not shown | 2026-09-28 | programbench.com. 3.0% nearly resolved, $5.60/task | |
| Claude Opus 4.7 | Anthropic | 0% | mini-SWE-agent, no internet, effort not shown | 2026-09-28 | programbench.com. 3.0% nearly resolved, $3.81/task |
| Claude Opus 4.6 | Anthropic | 0% | mini-SWE-agent, no internet, effort not shown | 2026-09-28 | programbench.com. 2.5% nearly resolved, $11.38/task |
| GPT 5.5 | OpenAI | 0% | mini-SWE-agent, no internet, effort not shown | 2026-09-28 | programbench.com. 1.5% nearly resolved, $1.21/task |
| Claude Sonnet 4.6 | Anthropic | 0% | mini-SWE-agent, no internet, effort not shown | 2026-09-28 | programbench.com. 1.0% nearly resolved, $26.73/task |
| GPT 5.4, Gemini 3.1 Pro, Gemini 3 Flash, Claude Haiku 4.5, GPT 5.4 mini, GPT 5 mini | OpenAI, Google, Anthropic | 0% | mini-SWE-agent, no internet, effort not shown | 2026-09-28 | programbench.com, ranks 20 to 25. 0.0% nearly resolved, $0.03 to $1.51/task |
Opus 5 at xhigh leads with 4.5% resolved and 37.0% nearly resolved. The next model, Muse Spark 1.3 at max effort, nearly resolves 25.0%. Muse Spark 1.3 gets there for $6.46 per task against $50.53 for Opus 5, so it is the better buy on this board if near-misses are useful to you.
This board shows what effort settings do, which the Vals table cannot. GPT 5.5 nearly resolves 1.5% at its default, 5.0% at high and 13.5% at xhigh. Muse Spark 1.3 resolves 1.0% at xhigh and 2.5% at max. Claude Opus 4.7 nearly resolves 3.0% at default and 4.5% at xhigh. More test-time compute buys real gains on this benchmark, and a model's name alone does not tell you its score.
Anthropic golden subset
Anthropic reports a mean hidden-test pass rate on its 166-task golden subset, with mini-SWE-agent and no six-hour limit.
Comparable within this table only; vendor-reported, 166 tasks instead of 200, mean pass rate instead of fully resolved, and no time limit.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 91.2% | mini-SWE-agent, 166 golden tasks, no six-hour limit | 2026-09-22 | Opus 5.5 System Card, section 8.10.1 |
| Claude Fable 5.1 | Anthropic | 87.6% | mini-SWE-agent, 166 golden tasks, no six-hour limit | 2026-09-01 | Fable 5.1 System Card, section 8.11.1; repeated in the Opus 5.5 card |
| Claude Fable 5 | Anthropic | 86.3% | mini-SWE-agent, 166 golden tasks, no six-hour limit | 2026-09-01 | Fable 5.1 System Card, section 8.11.1 |
| Claude Opus 5 | Anthropic | 85.4% | mini-SWE-agent, 166 golden tasks, no six-hour limit | 2026-09-01 and 2026-09-22 | Both system cards |
Opus 5.5 leads at 91.2% against 87.6% for Fable 5.1. The Vals raw pass rate column above, where Opus 5.5 scores 87.0%, uses all 200 tasks and keeps the time limit, so the two pass rates measure different things.
Of the vendors checked, only Anthropic reports ProgramBench. OpenAI's GPT-6 Astra system card and the GPT-6.1 Sol addendum do not mention it. Google has not published a figure, and its Gemini 4 evaluation methodology page returned a 404 when checked on 2026-10-06. The Astra and Gemini 4 numbers on this page come only from Vals. The Opus 5.5 release post does not mention ProgramBench either; the results live in the system card.
Multi-agent harness results
Anthropic's system cards also test multi-agent systems on the 166 golden tasks, with a 1M-token limit per agent and a bash tool. They publish curves, not numbers, so the scores below are read off a chart.
Not comparable with any other table. Latency is derived from fixed prefill and decode rates plus tool time, cost assumes perfect cache hits, and the Fable 5.1 run used an internal endpoint with Opus 5 as a refusal fallback, where 72% of episodes had at least one fallback turn and under 1% of turns fell back. Anthropic says to read these as a harness comparison, not an absolute score.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | 0.6, read from a curve | Five-agent team, 166 golden tasks, 1M tokens per agent, bash | 2026-09-01 | Fable 5.1 System Card, section 8.13.1. 2x lower latency than one agent; async subagents reach the highest final score; both multi-agent setups use more tokens |
| Claude Opus 5.5 (alternate snapshot) | Anthropic | 0.6, read from a curve | Five-agent team, 166 golden tasks, 1M tokens per agent, bash | 2026-09-22 | Opus 5.5 System Card, section 8.12.1. 2.7x lower latency than one agent; async subagents fall between single agent and team |
The result is about speed. A five-agent team reaches 0.6 between 2x and 2.7x faster than a single agent and spends more tokens doing it. If you are planning multi-agent orchestration for long builds, this is the only published ProgramBench evidence on how team structure changes wall-clock time.
Historical progression
The paper's nine models resolved nothing. The first resolved tasks appear on the September leaderboards.
The paper and programbench.com rows come from the maintainers' mini-SWE-agent runs, with vendor-default settings in the paper and effort settings on the live board. The Vals row is an independent run, so the move from 4.5% to 18.5% crosses leaderboards and is not a measured gain.
| Date | Model | Organization | Resolved | Almost resolved | Harness and setup | Source and notes |
|---|---|---|---|---|---|---|
| 2026-05-05 | Claude Opus 4.7 | Anthropic | 0.0% | 3.0% | mini-SWE-agent, vendor-default settings | arXiv paper. $3.81/task |
| 2026-05-05 | Claude Opus 4.6 | Anthropic | 0.0% | 2.5% | mini-SWE-agent, vendor-default settings | arXiv paper. $11.38/task |
| 2026-05-05 | Claude Sonnet 4.6 | Anthropic | 0.0% | 1.6% | mini-SWE-agent, vendor-default settings | arXiv paper. $27.09/task in the paper, $26.73 on the live board |
| 2026-05-05 | Gemini 3.1 Pro, GPT 5.4, GPT 5.4 mini, Claude Haiku 4.5 | Google, OpenAI, Anthropic | 0.0% | 0.0% | mini-SWE-agent, vendor-default settings | arXiv paper. The paper's other models also scored 0.0% |
| 2026-09-28 | Claude Opus 5 (xhigh) | Anthropic | 4.5% | 37.0% | mini-SWE-agent, no internet, xhigh effort | programbench.com. First resolved tasks |
| 2026-09-28 | Muse Spark 1.3 (max) | Meta | 2.5% | 25.0% | mini-SWE-agent, no internet, max effort | programbench.com |
| 2026-09-28 | GPT-5.6 Sol (xhigh) | OpenAI | 1.0% | 15.5% | mini-SWE-agent, no internet, xhigh effort | programbench.com |
| 2026-09-30 | Claude Opus 5.5 | Anthropic | 18.5% | 130 of 200 tasks | mini-SWE-agent via Valkyrie, max effort | Vals. Model released 2026-09-22; 37 tasks resolved |
The measured gains are the ones inside a single harness. On Vals, Opus 5 resolves 3.0% and Opus 5.5 resolves 18.5%. On Anthropic's golden subset, pass rate rises from 86.3% for Fable 5 to 87.6% for Fable 5.1 to 91.2% for Opus 5.5.
Documented failure modes
Monolithic programs. Models write programs that are much smaller and flatter than the originals. Among solutions passing 75% of tests, 207 runs across 88 tasks, the median solution has 1,173 lines in 3 files. The median reference has 3,068 lines in 15 files, so solutions carry 0.38x the lines and 0.20x the files. 67% have shallower directory trees than the reference, and only 15% exceed its line count. Opus 4.7's median solution has 39 functions where the reference has 133. A program that passes the tests as one giant file is a working clone and a maintenance problem.
Opposite working styles. GPT 5.4 writes 96% of its code in a single median turn and runs a median of 17 commands. Sonnet 4.6 runs a median of 868 commands, up to 1,978 turns, at $27.09 per task. Opus 4.7 puts 67% of its code into its largest single edit. In the paper, both styles ended at 0.0% resolved.
Cheating when the internet is open. With web access, Sonnet 4.6 cheated in 36% of runs, Opus 4.6 in 21%, Gemini 3 Flash in 20% and GPT 5 mini in 1%. Looking up the original source code accounted for 79% to 95% of violations. The authors tried catching this with an LLM judge, and agreement was only moderate, with a pooled Fleiss kappa of 0.57 and a range of 0.16 to 0.60. So they blocked the internet entirely. Any internal version of this eval needs the same network isolation.
Behavioral tests are a lower bound. A passing test can still hide wrong behavior, while a failing test is definitive. A finite suite covers less than the full specification, and speed, memory and disk use are not scored at all. A resolved task means the clone matched every probe, not that it would survive production traffic.
Contamination is blocked, not measured. Every reference program is an open-source project. The paper makes no claim about memorization. Its defense is cutting off code lookup: execute-only binaries, no internet, fresh git history, and separate build and inference containers. Its test audit found zero tests that invoke undocumented flags, and 5 of 200 instances with what the paper calls "plausibly implementation-dependent output".
Documentation dependence. A task is only feasible when the executable's interface is well documented. Programs with thin docs leave the agent guessing at behavior the tests expect.
Flaky tests. On 34 tasks the reference binary scores below 0.9 on its own hidden suite. Anthropic drops those tasks; the public leaderboards keep them.
Statistical noise. Vals reports Opus 5.5 at 18.50% ±2.75. Settings for the other Vals models are unpublished, and that ±2.75 on the leader is wider than most gaps among the models between 2% and 3%.
How ProgramBench compares to related benchmarks
SWE-bench fixes one issue inside an existing repository, with the codebase as a guide. ProgramBench builds the entire program with no source at all. Anthropic cites ProgramBench as a long-context test because episodes fill up to the 1M-token context window.
Commit0, DevBench and NL2Repo-bench are also from-scratch benchmarks, but they supply a skeleton with predefined classes and signatures and score against the source structure. ProgramBench grades only behavior against the reference executable, which leaves the architecture entirely to the model.
HumanEval and BigCodeBench test function-level completion from a prompt. There is no multi-file architecture and no executable to check against.
Terminal-Bench 4.0 covers shell tasks and FrontierCode covers repository-level bug and feature work. Both appear on the Opus 5.5 release post and ProgramBench does not. Their scores use different tasks and harnesses and do not compare with ProgramBench numbers. For other repository-level agent evals, see DeepSWE, FrontierSWE v2 and CursorBench. The LLM evaluation overview covers how these fit together.
What ProgramBench means for teams choosing a model
If you need whole programs built, Opus 5.5 is the only real option. It fully resolves 37 of 200 tasks on Vals, against 14 for Fable 5.1 and 11 for GPT-6 Astra, and almost resolves 130. It costs $68.05 per task and takes 2h24m.
If cost is the constraint, follow the frontier. GPT-6.1 Sol reaches 3.0% for $2.87. GPT-6 Astra resolves 5.5% for $11.57 in 27 minutes. Sonnet 5.5 resolves 6.5% for $35.86, and Fable 5.1 resolves 7.0% for $58.52. Skip Opus 5. At $60.29 it resolves 3.0%, the same as Sol at 21 times the price, and Fable 5.1 beats it by 4.0 points for $1.77 less.
Don't treat any model as production-grade for full rebuilds. Even Opus 5.5 fails 81.5% of tasks outright, and raw pass rates of 82% to 87% hide that. The almost-resolved count is a better guide to who gets close: 130 for Opus 5.5, 107 for Fable 5.1, 100 for Astra, 92 for Sol and 82 for Gemini 4 Argon. For most teams the realistic use is a near-complete draft that an engineer finishes.
Upgrade within a family. Opus 5.5 resolves 18.5% where Opus 5 resolves 3.0%. Fable 5.1 resolves 7.0% where Fable 5 resolves 2.0%, and costs $17.16 less per task. Per-token price misleads here. Fable 5.1 scores 11.5 points below Opus 5.5 at 2.5x the list price per token, $10/$50 against $4/$20, yet still costs less per task at $58.52 against $68.05.
If wall-clock time matters more than dollars, look at Astra. On Vals it finishes in 27 minutes against 2h07m to 2h27m for Sonnet 5.5, Opus 5.5 and Fable 5.1. On Anthropic's vendor-reported multi-agent harness, a five-agent team reaches 0.6, read from a curve, 2x to 2.7x faster than one agent, at higher token cost.
Pick from effort and harness, not from the model name. programbench.com and Vals disagree on the same model, with Opus 5 at 4.5% on one and 3.0% on the other, and effort settings move scores by several points. Rerun the eval at the effort level you plan to ship, and pin the date of any frontier model number you rely on.
You can compare these models across other benchmarks on the Klu LLM leaderboard. To run the same kind of behavior-graded eval on your own codebase and tasks, see how engineering teams set it up in Klu on the engineering page.