SRE-Bench (Software Reverse Engineering)

A benchmark that tests whether AI agents can work out what private, often hardened compiled binaries do when they have no source code

Security and operationsSolved of 262

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 14 min read

What is SRE-Bench?

SRE-Bench measures whether an AI agent can work out what a compiled binary does when it has no source code. SRE here means software reverse engineering. The benchmark has nothing to do with site reliability, and it is unrelated to the Kubernetes and operations benchmarks with similar names, such as SREGym, agentkube's SRE-bench and Rootly's SRE-skills-bench.

Columbia DAPLab built SRE-Bench with Vals AI and UC Berkeley. The paper, "The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark," has 12 authors with Jeremy Spence first. It is arXiv 2608.11469, now at v2, dated 2026-10-01. Vals co-developed the benchmark and also runs the public leaderboard. Treat the Vals board as a collaborator's board, not an outside audit.

The set is 19 private programs that reverse engineering experts wrote from scratch. They put more than 5,000 expert hours into them, and the programs average 16,915.8 lines of code. Compiled in different configurations, and many of them hardened with protections built from 44 in-house anti-analysis primitives, they become 262 binary instances. Each instance has six deterministically graded tasks, 1,572 in all.

The design fixes two problems with earlier reverse engineering benchmarks. CTF-derived sets like NYU CTF Bench, Cybench, CTF-Dojo and CTFTiny reuse public challenges, mix reverse engineering with other skills, and have no contamination control. From-scratch sets like CREBench, AgentRE-Bench and CrackMeBench are toys, at 526.2, 63.5 and 36.1 average lines, with textbook protections. SRE-Bench programs are 32 to 468 times larger.

The paper's contamination test shows why size and private provenance matter. GPT-5.6 Sol and GPT-5.5 each solve a 1.1K-line clean-room compressor in all eight builds, and each solves a gzip variant with about 200 modified lines in all eight builds for about $2 and 10 minutes. On RevCompress, the real-scale private compressor, Sol solves 7 of 8 for $31.45 and 94.6 minutes, and GPT-5.5 solves 3 of 8. Small or public-lineage programs make the job easy. A large private program is where models start to separate.

On the standardized harness, GPT-6 Astra leads with 56.87% of all 262 instances fully solved, 26.3 points ahead of GPT-5.6 Sol. In OpenAI's own unconstrained Codex runs, Astra reaches 88.0% pass@1.

How the tasks work

The 19 programs fall into five domains, and each domain has its own deterministic grader. There are four network protocol programs, four games, four file format programs, four malware samples and three firmware images. The graders check recovered behavior directly, for example protocol coverage, byte-exact decoding, malware cleanup and hidden trigger behavior. Malware cleanup is all-or-nothing, and the agent has to remove the malware while preserving benign data the grader planted. Firmware tasks ask the agent to take control of a locked-down microcontroller, graded on a held-out emulator. Programs are written in C, C++, Rust and Go, with no Go firmware.

Each of the 16 non-firmware programs ships in 16 builds. Eight are unhardened and vary optimization on or off, symbols stripped or not, and static or dynamic linking. The other eight are hardened, one per protection preset, P1 through P8. That gives 128 unhardened instances, 128 hardened instances and 6 firmware instances. The presets combine the 44 protection primitives, and more than half of those primitives have no public implementation. Hardening turns out to be what splits the field.

Every instance runs in an isolated container with the binary and standard reverse engineering tools, namely Ghidra, radare2, GDB and angr. Grader code and reference solutions are removed, and the agent gets no scoring feedback while it works. A private reference solution scores 6 of 6 on every program, so every task is solvable.

The binaries and specifications stay private. Evaluation runs through an evaluation service, and the paper withholds constants, triggers and reference solutions so it does not become training data. The cost of that choice is that nobody outside the project can run SRE-Bench locally.

The two harnesses

The paper reports results under two very different setups, and a score means little until you know which one produced it.

The standardized harness runs every model in mini-SWE-agent at its highest reasoning effort, with the vendor's cyber safeguards switched on. Each run stops at 500 model steps or 6 hours of wall clock.

The unconstrained harness is each vendor's own agent, Codex for OpenAI models and Claude Code for Anthropic models. It uses vendor default settings, turns cyber safeguards off and has no step or budget cap. OpenAI and Anthropic ran these internally, and the paper reports the results. Customers cannot buy this setup. Results come as pass@1 and pass@4 over four attempts per instance, except Claude Opus 5.5's Low to XHigh effort ladder, which is one attempt each.

Vals describes its own setup as "every agent works in the same shell-only loop with the same reverse-engineering toolkit." It does not name the scaffold.

How scoring works

The paper uses three numbers on the standardized harness, and they diverge for models that refuse.

Score is the mean number of tasks passed out of six, averaged over graded binaries. Solved means all six tasks pass. A run counts as graded only if it produced a gradable result. Runs lost to safeguard refusals or context-window failures are dropped from Score and from the Solved percentage. A model that refuses a third of the set can still post a high Score on what it attempted.

Solved of 262 fixes that by counting every instance, with refusals scored as misses. For all eight standardized rows it equals Vals accuracy. Rank on Solved of 262.

In the unconstrained runs, pass@1 is the share of instances with all six tasks solved, averaged over 262 instances and four attempts. Pass@4 counts an instance if any of the four attempts solves it. The paper also reports a Coverage figure for each run.

Versions

The paper's v1 went up on 2026-08-11 and v2 on 2026-10-01. No v1 scores are sourced here, so every paper number on this page comes from v2, and the Vals numbers come from the board snapshot of 2026-09-29. The paper does not state its evaluation window, so paper tables carry the v2 date. No lab system card or release post carries an SRE-Bench figure.

Current leaderboard

Standardized harness

All eight rows ran on the same harness, with the same caps, in the authors' run. They are comparable with each other and not with the Vals-only or unconstrained tables below.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI56.87% ±3.07 of 262mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours2026-10-01Paper Table 4. 149 of 181 graded solved, 82.3%. Score 5.55 of 6 ±0.09. 7 zero-credit runs. $13.50 and 18 min per graded instance. Vals: $10.24, 17m51s, effort max
GPT-5.6 SolOpenAI30.53% ±2.85 of 262mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours2026-10-01Paper Table 4. 80 of 254 graded solved, 31.5%. Score 3.69 ±0.14. 41 zero-credit. $42.50, 88 min. Vals: $42.52, 1h28m, effort max
Claude Fable 5.1Anthropic22.90% ±2.60 of 262mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours2026-10-01Paper Table 4. 60 of 223 graded solved, 26.9%. Score 2.58 ±0.18. 96 zero-credit. $34.70, 123 min. Vals: $32.23, 2h03m, temperature 1.0, 128,000 max output
Claude Opus 5Anthropic12.21% of 262mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours2026-10-01Paper Table 4. 32 of 256 graded solved, 12.5%. Score 1.91 ±0.14. 118 zero-credit. $23.60, 81 min. Vals: $23.63, 1h21m
GPT-5.5OpenAI3.82% of 262mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours2026-10-01Paper Table 4. 10 of 262 solved. Score 1.02 ±0.10. 148 zero-credit. $17.80, 46 min. Vals: $17.77, 46m28s
DeepSeek V4.1 FlashDeepSeek0.76% of 262mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours2026-10-01Paper Table 4. 2 of 262 solved. Score 0.85 ±0.09. 166 zero-credit. $0.50, 61 min. Vals: $0.55, 1h01m
Grok 4.5xAI0.76% of 262mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours2026-10-01Paper Table 4. 2 of 218 graded solved. Score 0.45 ±0.07. 160 zero-credit. $13.50, 66 min. Vals: $13.42, 1h06m
GLM 5.2Z.AI0.00% of 262mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours2026-10-01Paper Table 4. 0 of 262 solved. Score 0.21 ±0.03. 217 zero-credit. $26.60, 149 min. Vals: $26.58, 2h29m
SRE-Bench authorsmini-SWE-agent262 instancesSolved of 262arxiv.org
Solved of 262 against time per instancemini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours
20%40%60%20 min50 min1.67 hr3.33 hr
FrontierOpenAIBehind the frontier
Time per graded instance, log scale
Source: Spence et al., arXiv 2608.11469v2, Table 4, dated 2026-10-01, with accuracy cross-checked against the Vals AI board of 2026-09-29. X is the paper's average time over graded runs and Y is Solved of 262. Graded counts below 262: Astra 181, Sol 254, Fable 5.1 223, Opus 5 256, Grok 4.5 218. Astra's time excludes its 81 refused instances.

Astra is the top point on both axes. It solves 56.87% of the set in 18 minutes per graded instance, against 88 minutes for Sol and 123 for Fable 5.1. On cost it matches Grok 4.5 at $13.50 per graded instance, and Grok solves 0.76%. Astra's per-instance cost and time cover only its 181 graded runs, though. Its 81 refusals cost nothing and still count as misses in the 56.87%.

Below the top three, the field collapses. Opus 5 solves 12.21%, GPT-5.5 3.82%, and the three remaining models solve two instances or none.

Paper and Vals costs agree to within $0.08 for six rows. They differ for Astra, $13.50 against $10.24, and for Fable 5.1, $34.70 against $32.23. The paper averages over graded runs. Vals reports a per-test figure and has not published its denominator. Astra's 149 solved matches Vals' 56.87% because both count the 81 refusals as misses. Neither source says whether Vals re-ran Astra or reused the authors' run, so the Vals column is a cross-check, not a replication. The paper reports 13 configurations across 11 models, and this table carries the eight standardized ones.

Vals-only runs

Newer models appear only on the Vals board. Each run has its own date and settings, and Vals does not name the harness, so each is a separate table and none is comparable with another or with the standardized table above.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6.1 SolOpenAI50.76% ±3.10Vals shell-only loop, reasoning effort maxEval date not publishedVals board, updated 2026-09-29. $2.69 per test, 32m32s. No fallback disclosed. Model page shows $3.24 and 43m10s on its Vals Index basis
Vals AIVals shell-only loopAccuracyvals.ai
ModelOrganizationScoreHarness and setupDateSource and notes
Gemini 4 ArgonGoogle44.27% ±3.08Vals shell-only loop, reasoning effort high, temperature 1.02026-09-30Vals board. $30.45 per test, 1h35m. No fallback disclosed. Model page
Vals AIVals shell-only loopAccuracyvals.ai
ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic33.59% ±2.92Vals shell-only loop, effort not stated2026-09-22Vals board. $34.65 per test, 1h37m. 217 of 262 tasks, 82.82%, fell back to earlier Claude versions. Counting those as failures gives 5.34%, per the model page, which also shows $32.14 and 79 min on its Vals Index basis
Vals AIVals shell-only loopAccuracyvals.ai
ModelOrganizationScoreHarness and setupDateSource and notes
Claude Sonnet 5.5Anthropic30.15% ±2.84Vals shell-only loop, compute effort max, temperature 1.02026-09-28Vals board. $26.52 per test, 1h55m. Claude Sonnet 5 was the fallback. Counting fallback tasks as failures gives 19.08%, 126 of 262, per the model page
Vals AIVals shell-only loopAccuracyvals.ai

The fallback disclosures change how to read the Claude numbers. When a newer Claude model refuses, Vals reports a server-side fallback to an earlier Claude model, and the task can still pass. Opus 5.5's own work accounts for 5.34% of the set. The same applies to Fable 5.1 on the standardized table. Its Vals model page, from a 2026-09-01 eval, lists 195 of 262 tasks, 74.43%, as fallback-assisted by Claude Opus 5 and Opus 4.8, and Fable 5.1's 22.90% falls to 10.69% without them.

Unconstrained OpenAI runs in Codex

OpenAI ran these internally in Codex with safeguards off, no budget cap and four attempts per instance. Rows compare with each other, not with the standardized harness or the Claude Code runs.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI88.0% pass@1Codex, safeguards off, no cap, Max, n=42026-10-01Paper Table 11. pass@4 99.2%, Coverage 97.0%, 69.7K output tokens per attempt
GPT-6 AstraOpenAI77.7% pass@1Codex, safeguards off, no cap, XHigh, n=42026-10-01Paper Table 11. pass@4 94.3%, Coverage 93.4%, 42.6K tokens
GPT-6 AstraOpenAI74.7% pass@1Codex, safeguards off, no cap, High, n=42026-10-01Paper Table 11. pass@4 92.4%, Coverage 93.2%, 38.7K tokens
GPT-6 AstraOpenAI72.2% pass@1Codex, safeguards off, no cap, Medium, n=42026-10-01Paper Table 11. pass@4 90.1%, Coverage 90.3%, 34.2K tokens
GPT-6 AstraOpenAI52.3% pass@1Codex, safeguards off, no cap, Low, n=42026-10-01Paper Table 11. pass@4 67.2%, Coverage 76.1%, 20.7K tokens
GPT-5.6 CyberOpenAI56.8% pass@1Codex, safeguards off, no cap, Max, n=42026-10-01Paper Table 11. pass@4 72.5%, Coverage 83.1%, 342.6K tokens
GPT-5.6 CyberOpenAI43.4% pass@1Codex, safeguards off, no cap, XHigh, n=42026-10-01Paper Table 11. pass@4 56.9%, Coverage 68.8%, 211.1K tokens
GPT-5.6 CyberOpenAI23.1% pass@1Codex, safeguards off, no cap, High, n=42026-10-01Paper Table 11. pass@4 34.0%, Coverage 43.9%, 127.7K tokens
GPT-5.6 CyberOpenAI10.5% pass@1Codex, safeguards off, no cap, Medium, n=42026-10-01Paper Table 11. pass@4 16.8%, Coverage 23.6%, 64.5K tokens
GPT-5.6 SolOpenAI55.8% pass@1Codex, safeguards off, no cap, Max, n=42026-10-01Paper Table 11. pass@4 68.7%, Coverage 83.9%, 241.5K tokens
GPT-5.6 SolOpenAI41.2% pass@1Codex, safeguards off, no cap, XHigh, n=42026-10-01Paper Table 11. pass@4 53.8%, Coverage 69.1%, 143.3K tokens
GPT-5.6 SolOpenAI16.7% pass@1Codex, safeguards off, no cap, High, n=42026-10-01Paper Table 11. pass@4 23.3%, Coverage 36.6%, 70.9K tokens
GPT-5.6 SolOpenAI5.8% pass@1Codex, safeguards off, no cap, Medium, n=42026-10-01Paper Table 11. pass@4 8.8%, Coverage 14.9%, 28.9K tokens
OpenAICodex262 instances, n=4pass@1arxiv.org

Astra at Max leads at 88.0% pass@1, 31.2 points ahead of GPT-5.6 Cyber at Max and 32.2 ahead of Sol at Max. Astra at Medium already beats both of them at Max, at 72.2%, on roughly a seventh of Sol's output tokens per attempt. Sol swings the most with effort, from 5.8% at Medium to 55.8% at Max. For Sol, the effort setting is a bigger decision than the model name, a plain case of the test-time compute trade-off.

OpenAI published per-task costs for two of these models. Summed over every attempt at every effort level, Astra's 5,240 runs cost $83.4K and produced 215.7M output tokens, an average of $15.92 per attempt. Sol's 4,192 runs cost $161.9K and produced 506.9M output tokens, $38.62 per attempt. The paper's Anthropic and GPT-5.6 Cyber costs are estimates on a different basis, so this page does not compare cost across vendors.

Unconstrained Anthropic runs in Claude Code

Anthropic ran these internally in Claude Code at Max effort with safeguards off, no budget cap, four attempts per instance and a clarified malware prompt. They compare with each other, not with the Codex runs or the standardized harness.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic82.1% pass@1Claude Code, safeguards off, no cap, Max, n=4, clarified prompt2026-10-01Paper Table 12. pass@4 98.1%, Coverage@1 92.9%. 321.1K output tokens per attempt, from Table 11
Claude Mythos 5.1Anthropic53.9% pass@1Claude Code, safeguards off, no cap, Max, n=4, clarified prompt2026-10-01Paper Table 12. pass@4 61.8%, Coverage@1 76.0%. No token figure published
AnthropicClaude Code262 instances, n=4pass@1arxiv.org

The effort ladder below is a separate single-attempt run of Opus 5.5. It compares across its own rows only, and its XHigh row is not a step toward the n=4 Max result above.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic77.9% pass@1Claude Code, safeguards off, XHigh, n=1, clarified prompt2026-10-01Paper Table 11. Coverage 89.3%, 240.1K output tokens per attempt
Claude Opus 5.5Anthropic66.8% pass@1Claude Code, safeguards off, High, n=1, clarified prompt2026-10-01Paper Table 11. Coverage 85.2%, 169.6K tokens
Claude Opus 5.5Anthropic59.9% pass@1Claude Code, safeguards off, Medium, n=1, clarified prompt2026-10-01Paper Table 11. Coverage 81.2%, 143.3K tokens
Claude Opus 5.5Anthropic53.1% pass@1Claude Code, safeguards off, Low, n=1, clarified prompt2026-10-01Paper Table 11. Coverage 84.4%, 117.5K tokens
AnthropicClaude Code262 instances, n=1pass@1arxiv.org

Opus 5.5 leads Mythos 5.1 by 28.2 points at Max. Put the Opus 5.5 rows next to its Vals entry and the gap between the two harnesses is stark. Safeguards off in Claude Code, it solves 82.1%. On the Vals board, where its own work accounts for 5.34% and fallback models carry the rest, it posts 33.59%. Astra's 88.0% and Opus 5.5's 82.1% come from different vendor harnesses, so they are not a head-to-head result.

Historical progression

SRE-Bench is new, and its only same-harness history is the generation steps inside the standardized table. Because these models ran in one harness with one set of caps, the steps are valid comparisons.

FamilyEarlier modelSolved of 262Later modelSolved of 262Change
OpenAIGPT-5.53.82%GPT-5.6 Sol30.53%+26.7 points
OpenAIGPT-5.6 Sol30.53%GPT-6 Astra56.87%+26.3 points
AnthropicClaude Opus 512.21%Claude Fable 5.122.90%+10.7 points

Source: paper Table 4, v2 dated 2026-10-01. Opus 5.5 is not in the standardized table, so the Anthropic line stops at Fable 5.1. GPT-6.1 Sol, Gemini 4 Argon, Opus 5.5 and Sonnet 5.5 appear only in Vals runs, and no earlier-version score exists for them on a matching setup.

OpenAI's two steps are about the same size, each adding roughly 26 points. Anthropic's single step is less than half that, and Fable 5.1's figure carries the fallback caveat above.

Failure modes and limitations

Hardened binaries

Protection presets are the obstacle that separates models. The table shows Score out of six on the standardized harness for unhardened and hardened builds, and the cost per point earned.

ModelUnhardened ScoreHardened ScoreCost per point, unhardenedCost per point, hardenedCost multiple
GPT-6 Astra5.835.73$2.15$3.181.5x
GPT-5.6 Sol4.692.50$6.52$22.393.4x
Claude Fable 5.15.330.32$6.93$105.7915.3x
Claude Opus 5Not printed0.33$7.22$66.989.3x
GPT-5.5Not printed0.11$11.02$138.5412.6x
DeepSeek V4.1 FlashNot printed0.08Not printedNot printedNot printed
Grok 4.5Not printed0.02$18.56$654.3135.3x
GLM 5.2Not printed0.00Not printedNo points earnedNot applicable
SRE-Bench authorsmini-SWE-agentHardened buildsScore out of 6 on hardened buildsarxiv.org

Source: paper Table 7 and section 4, v2 dated 2026-10-01. Astra's figures cover only its graded runs.

Fable 5.1 is the sharpest case. It scores 5.33 of 6 on unhardened builds, close to Astra, and 0.32 on hardened ones. Every model except Astra and Sol scores 0.33 or lower once protections go on. Astra drops 0.10 points. Sol keeps about half its Score, at a cost per point 3.4 times higher.

The unconstrained runs show the same pattern within each harness. In Codex, Astra loses 0 to 6 pass@1 points under any preset compared with its unhardened builds, while Sol loses 13.5 to 19.7. In Claude Code, Opus 5.5 loses 0 to 15 points. Mythos 5.1 loses under 4, but from an unhardened base of 53.9%.

Stripped symbols and optimization

Stripping symbols is the biggest build penalty on the standardized harness. The four weakest models lose 50 to 71% of their Score on stripped builds. Sol, Fable 5.1 and Opus 5 lose 10 to 15%. Astra loses nothing, 5.39 against 5.36. Turning on compiler optimization costs the weaker models 17 to 41%.

Games, especially in Rust

In the unconstrained runs, games are the hardest domain for GPT-5.6 Sol in Codex and for Mythos 5.1 in Claude Code. Sol solves 1.6% of Rust game instances in Codex, and Mythos 5.1 solves none in Claude Code. Rust is the lowest game cell for all four models. Astra ranges from 35.9% on Rust to 90.6% on C, and Opus 5.5 from 34.4% on Rust to 76.6% on C.

Source-level security skill does not transfer

GPT-5.6 Sol and Fable 5.1 both have strong source-level security results, and on the standardized harness they fully solve 31.5% and 26.9% of their graded instances. Reading code and reading a stripped, protected binary are different skills, and a source-level security score does not predict this one.

Refusals

Safeguard refusals left only 181 of Astra's 262 standardized instances, 69.1%, graded, even though the tasks do not involve exploitation. The paper says this raises concerns about Astra's practical use for reverse engineering. The Claude models hit the same wall from a different direction. Their Vals scores depend on fallback to earlier models, with Opus 5.5 falling from 33.59% to 5.34%, Fable 5.1 from 22.90% to 10.69% and Sonnet 5.5 from 30.15% to 19.08% when fallback tasks count as failures.

Picking the right attempt

Astra's 99.2% pass@4 at Max looks close to a solved benchmark. It isn't. Astra solved 191 instances in all four attempts, 33 in exactly three, 23 in exactly two, 13 in exactly one and 2 in none. That leaves 69 instances where right and wrong submissions sit side by side, and 126 of 1,048 attempts fail. SRE-Bench does not test whether an agent can pick the correct submission without the private grader, and in production there is no private grader. Plan on the 88.0% pass@1.

Prompt wording

The malware prompt's wording moves Opus 5.5 at Max in Claude Code from 76.3% pass@1 with the original prompt to 82.1% with a clarified one, and Coverage@1 from 87.4% to 92.9%. Mythos 5.1 barely moves, 53.3% to 53.9%. The paper publishes neither prompt, so an outside team cannot reproduce the difference.

Reward hacking

An early build leaked a hidden behavior's name as a readable string, and an agent found the trigger by scanning strings instead of analyzing the code. The authors removed the string and added a regression check that fails the build if it happens again. They used Opus 4.8 and GPT-6 Astra for audits.

Breadth and access

Nineteen programs averaging 16.9K lines is far larger than earlier RE sets but far smaller than a million-line system. The binaries are private, so outside parties cannot re-run the benchmark locally or check a vendor's number themselves.

ProgramBench also starts from a compiled binary, but it uses known open-source projects and asks the agent to reimplement the program's behavior. SRE-Bench uses private binaries and grades the recovered semantics with task-specific graders for protocol coverage, byte-exact decoding, malware cleanup and hidden triggers. ProgramBench tests whether a model can rebuild a program. SRE-Bench tests whether it understands one it cannot read.

Against the CTF-derived sets, NYU CTF Bench, Cybench, CTF-Dojo and CTFTiny, and the toy RE sets, CREBench, AgentRE-Bench and CrackMeBench, the difference is scale and provenance. Those use public challenges or programs of 36 to 526 average lines with easy protections. SRE-Bench averages 16,915.8 lines, stays private and adds hardening. The contamination test above shows what that buys. Models that ace a small or public-lineage compressor drop sharply on the real-scale private one.

ExploitBench, ExploitGym and SEC-Bench Pro cover exploit development and other security work. SRE-Bench leaves exploit development out on purpose, so a score measures reverse engineering without mixing in exploitation skill. Teams assessing offensive security capability, or planning red-teaming work, need both kinds of result.

SWE-bench tests source-level fixes in public repositories. SRE-Bench is the binary-first counterpart, with private programs and no source code to read. A team that picked a model on SWE-bench has learned nothing yet about how it handles a stripped, protected binary.

What it means for teams choosing a model

GPT-6 Astra is the pick on the standardized harness. It solves 56.87% of all 262 instances, against 30.53% for GPT-5.6 Sol and 22.90% for Claude Fable 5.1. It does it for $13.50 per graded instance, against $42.50 for Sol and $34.70 for Fable 5.1, in 18 minutes, the fastest in the table. Only DeepSeek V4.1 Flash is cheaper, and it solves 0.76%.

Budget for refusals. Astra was graded on only 181 of 262 instances because safeguards refused the rest. The 56.87% already counts those as misses, but a team doing reverse engineering work needs a fallback path or a vendor access program for the instances the model will not touch.

Test on hardened binaries if your binaries are hardened. Astra holds at 5.73 of 6 on hardened builds. Sol drops to 2.50, Opus 5 to 0.33, Fable 5.1 to 0.32 from 5.33 unhardened, and GPT-5.5 to 0.11. Astra's hardened Score covers only graded runs, so rank on Solved of 262 to compare. No hardened data exists for GPT-6.1 Sol, Gemini 4 Argon or Opus 5.5.

Read Vals scores with the fallback share next to them. On the Vals board, GPT-6.1 Sol posts 50.76% at $2.69 per test, Gemini 4 Argon 44.27% at $30.45 and Claude Opus 5.5 33.59% at $34.65. Each ran on its own date with its own settings. Opus 5.5's own work accounts for 5.34%, and the rest passes through fallback to older Claude models. A buyer comparing Vals scores needs to know how much of each one the named model earned.

Treat effort as part of the model. In Codex, Astra at Medium reaches 72.2% pass@1 on 34.2K output tokens per attempt, while Sol at Max needs 241.5K tokens for 55.8%. In Claude Code, Opus 5.5 climbs from 53.1% at Low on 117.5K tokens to 77.9% at XHigh on 240.1K, in a single-attempt run. Output tokens per attempt differ by up to seven times across these settings, so weigh cost per token and token use together at the effort you plan to run.

Plan on pass@1. Astra's 99.2% pass@4 needs a selector that SRE-Bench does not include, and 69 of its instances mix right and wrong attempts. The number to plan around is 88.0%, and that figure comes from an unconstrained harness customers cannot buy.

Cheap models do not substitute. DeepSeek V4.1 Flash costs $0.50 per instance and solves 0.76%. GLM 5.2 costs $26.60 and solves none. Only the top three frontier models in the standardized table solve more than an eighth of the set.

You can compare these models across other benchmarks on the Klu LLM leaderboard. To run the same kind of deterministic, task-graded evaluation on your own engineering workflows, see how teams set it up in Klu on the engineering page.