What is SRE-Bench?
SRE-Bench measures whether an AI agent can work out what a compiled binary does when it has no source code. SRE here means software reverse engineering. The benchmark has nothing to do with site reliability, and it is unrelated to the Kubernetes and operations benchmarks with similar names, such as SREGym, agentkube's SRE-bench and Rootly's SRE-skills-bench.
Columbia DAPLab built SRE-Bench with Vals AI and UC Berkeley. The paper, "The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark," has 12 authors with Jeremy Spence first. It is arXiv 2608.11469, now at v2, dated 2026-10-01. Vals co-developed the benchmark and also runs the public leaderboard. Treat the Vals board as a collaborator's board, not an outside audit.
The set is 19 private programs that reverse engineering experts wrote from scratch. They put more than 5,000 expert hours into them, and the programs average 16,915.8 lines of code. Compiled in different configurations, and many of them hardened with protections built from 44 in-house anti-analysis primitives, they become 262 binary instances. Each instance has six deterministically graded tasks, 1,572 in all.
The design fixes two problems with earlier reverse engineering benchmarks. CTF-derived sets like NYU CTF Bench, Cybench, CTF-Dojo and CTFTiny reuse public challenges, mix reverse engineering with other skills, and have no contamination control. From-scratch sets like CREBench, AgentRE-Bench and CrackMeBench are toys, at 526.2, 63.5 and 36.1 average lines, with textbook protections. SRE-Bench programs are 32 to 468 times larger.
The paper's contamination test shows why size and private provenance matter. GPT-5.6 Sol and GPT-5.5 each solve a 1.1K-line clean-room compressor in all eight builds, and each solves a gzip variant with about 200 modified lines in all eight builds for about $2 and 10 minutes. On RevCompress, the real-scale private compressor, Sol solves 7 of 8 for $31.45 and 94.6 minutes, and GPT-5.5 solves 3 of 8. Small or public-lineage programs make the job easy. A large private program is where models start to separate.
On the standardized harness, GPT-6 Astra leads with 56.87% of all 262 instances fully solved, 26.3 points ahead of GPT-5.6 Sol. In OpenAI's own unconstrained Codex runs, Astra reaches 88.0% pass@1.
How the tasks work
The 19 programs fall into five domains, and each domain has its own deterministic grader. There are four network protocol programs, four games, four file format programs, four malware samples and three firmware images. The graders check recovered behavior directly, for example protocol coverage, byte-exact decoding, malware cleanup and hidden trigger behavior. Malware cleanup is all-or-nothing, and the agent has to remove the malware while preserving benign data the grader planted. Firmware tasks ask the agent to take control of a locked-down microcontroller, graded on a held-out emulator. Programs are written in C, C++, Rust and Go, with no Go firmware.
Each of the 16 non-firmware programs ships in 16 builds. Eight are unhardened and vary optimization on or off, symbols stripped or not, and static or dynamic linking. The other eight are hardened, one per protection preset, P1 through P8. That gives 128 unhardened instances, 128 hardened instances and 6 firmware instances. The presets combine the 44 protection primitives, and more than half of those primitives have no public implementation. Hardening turns out to be what splits the field.
Every instance runs in an isolated container with the binary and standard reverse engineering tools, namely Ghidra, radare2, GDB and angr. Grader code and reference solutions are removed, and the agent gets no scoring feedback while it works. A private reference solution scores 6 of 6 on every program, so every task is solvable.
The binaries and specifications stay private. Evaluation runs through an evaluation service, and the paper withholds constants, triggers and reference solutions so it does not become training data. The cost of that choice is that nobody outside the project can run SRE-Bench locally.
The two harnesses
The paper reports results under two very different setups, and a score means little until you know which one produced it.
The standardized harness runs every model in mini-SWE-agent at its highest reasoning effort, with the vendor's cyber safeguards switched on. Each run stops at 500 model steps or 6 hours of wall clock.
The unconstrained harness is each vendor's own agent, Codex for OpenAI models and Claude Code for Anthropic models. It uses vendor default settings, turns cyber safeguards off and has no step or budget cap. OpenAI and Anthropic ran these internally, and the paper reports the results. Customers cannot buy this setup. Results come as pass@1 and pass@4 over four attempts per instance, except Claude Opus 5.5's Low to XHigh effort ladder, which is one attempt each.
Vals describes its own setup as "every agent works in the same shell-only loop with the same reverse-engineering toolkit." It does not name the scaffold.
How scoring works
The paper uses three numbers on the standardized harness, and they diverge for models that refuse.
Score is the mean number of tasks passed out of six, averaged over graded binaries. Solved means all six tasks pass. A run counts as graded only if it produced a gradable result. Runs lost to safeguard refusals or context-window failures are dropped from Score and from the Solved percentage. A model that refuses a third of the set can still post a high Score on what it attempted.
Solved of 262 fixes that by counting every instance, with refusals scored as misses. For all eight standardized rows it equals Vals accuracy. Rank on Solved of 262.
In the unconstrained runs, pass@1 is the share of instances with all six tasks solved, averaged over 262 instances and four attempts. Pass@4 counts an instance if any of the four attempts solves it. The paper also reports a Coverage figure for each run.
Versions
The paper's v1 went up on 2026-08-11 and v2 on 2026-10-01. No v1 scores are sourced here, so every paper number on this page comes from v2, and the Vals numbers come from the board snapshot of 2026-09-29. The paper does not state its evaluation window, so paper tables carry the v2 date. No lab system card or release post carries an SRE-Bench figure.
Current leaderboard
Standardized harness
All eight rows ran on the same harness, with the same caps, in the authors' run. They are comparable with each other and not with the Vals-only or unconstrained tables below.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 56.87% ±3.07 of 262 | mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours | 2026-10-01 | Paper Table 4. 149 of 181 graded solved, 82.3%. Score 5.55 of 6 ±0.09. 7 zero-credit runs. $13.50 and 18 min per graded instance. Vals: $10.24, 17m51s, effort max |
| GPT-5.6 Sol | OpenAI | 30.53% ±2.85 of 262 | mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours | 2026-10-01 | Paper Table 4. 80 of 254 graded solved, 31.5%. Score 3.69 ±0.14. 41 zero-credit. $42.50, 88 min. Vals: $42.52, 1h28m, effort max |
| Claude Fable 5.1 | Anthropic | 22.90% ±2.60 of 262 | mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours | 2026-10-01 | Paper Table 4. 60 of 223 graded solved, 26.9%. Score 2.58 ±0.18. 96 zero-credit. $34.70, 123 min. Vals: $32.23, 2h03m, temperature 1.0, 128,000 max output |
| Claude Opus 5 | Anthropic | 12.21% of 262 | mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours | 2026-10-01 | Paper Table 4. 32 of 256 graded solved, 12.5%. Score 1.91 ±0.14. 118 zero-credit. $23.60, 81 min. Vals: $23.63, 1h21m |
| GPT-5.5 | OpenAI | 3.82% of 262 | mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours | 2026-10-01 | Paper Table 4. 10 of 262 solved. Score 1.02 ±0.10. 148 zero-credit. $17.80, 46 min. Vals: $17.77, 46m28s |
| DeepSeek V4.1 Flash | DeepSeek | 0.76% of 262 | mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours | 2026-10-01 | Paper Table 4. 2 of 262 solved. Score 0.85 ±0.09. 166 zero-credit. $0.50, 61 min. Vals: $0.55, 1h01m |
| Grok 4.5 | xAI | 0.76% of 262 | mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours | 2026-10-01 | Paper Table 4. 2 of 218 graded solved. Score 0.45 ±0.07. 160 zero-credit. $13.50, 66 min. Vals: $13.42, 1h06m |
| GLM 5.2 | Z.AI | 0.00% of 262 | mini-SWE-agent, safeguards on, max effort, 500 steps, 6 hours | 2026-10-01 | Paper Table 4. 0 of 262 solved. Score 0.21 ±0.03. 217 zero-credit. $26.60, 149 min. Vals: $26.58, 2h29m |
Astra is the top point on both axes. It solves 56.87% of the set in 18 minutes per graded instance, against 88 minutes for Sol and 123 for Fable 5.1. On cost it matches Grok 4.5 at $13.50 per graded instance, and Grok solves 0.76%. Astra's per-instance cost and time cover only its 181 graded runs, though. Its 81 refusals cost nothing and still count as misses in the 56.87%.
Below the top three, the field collapses. Opus 5 solves 12.21%, GPT-5.5 3.82%, and the three remaining models solve two instances or none.
Paper and Vals costs agree to within $0.08 for six rows. They differ for Astra, $13.50 against $10.24, and for Fable 5.1, $34.70 against $32.23. The paper averages over graded runs. Vals reports a per-test figure and has not published its denominator. Astra's 149 solved matches Vals' 56.87% because both count the 81 refusals as misses. Neither source says whether Vals re-ran Astra or reused the authors' run, so the Vals column is a cross-check, not a replication. The paper reports 13 configurations across 11 models, and this table carries the eight standardized ones.
Vals-only runs
Newer models appear only on the Vals board. Each run has its own date and settings, and Vals does not name the harness, so each is a separate table and none is comparable with another or with the standardized table above.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-6.1 Sol | OpenAI | 50.76% ±3.10 | Vals shell-only loop, reasoning effort max | Eval date not published | Vals board, updated 2026-09-29. $2.69 per test, 32m32s. No fallback disclosed. Model page shows $3.24 and 43m10s on its Vals Index basis |
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Gemini 4 Argon | 44.27% ±3.08 | Vals shell-only loop, reasoning effort high, temperature 1.0 | 2026-09-30 | Vals board. $30.45 per test, 1h35m. No fallback disclosed. Model page |
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 33.59% ±2.92 | Vals shell-only loop, effort not stated | 2026-09-22 | Vals board. $34.65 per test, 1h37m. 217 of 262 tasks, 82.82%, fell back to earlier Claude versions. Counting those as failures gives 5.34%, per the model page, which also shows $32.14 and 79 min on its Vals Index basis |
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Sonnet 5.5 | Anthropic | 30.15% ±2.84 | Vals shell-only loop, compute effort max, temperature 1.0 | 2026-09-28 | Vals board. $26.52 per test, 1h55m. Claude Sonnet 5 was the fallback. Counting fallback tasks as failures gives 19.08%, 126 of 262, per the model page |
The fallback disclosures change how to read the Claude numbers. When a newer Claude model refuses, Vals reports a server-side fallback to an earlier Claude model, and the task can still pass. Opus 5.5's own work accounts for 5.34% of the set. The same applies to Fable 5.1 on the standardized table. Its Vals model page, from a 2026-09-01 eval, lists 195 of 262 tasks, 74.43%, as fallback-assisted by Claude Opus 5 and Opus 4.8, and Fable 5.1's 22.90% falls to 10.69% without them.
Unconstrained OpenAI runs in Codex
OpenAI ran these internally in Codex with safeguards off, no budget cap and four attempts per instance. Rows compare with each other, not with the standardized harness or the Claude Code runs.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 88.0% pass@1 | Codex, safeguards off, no cap, Max, n=4 | 2026-10-01 | Paper Table 11. pass@4 99.2%, Coverage 97.0%, 69.7K output tokens per attempt |
| GPT-6 Astra | OpenAI | 77.7% pass@1 | Codex, safeguards off, no cap, XHigh, n=4 | 2026-10-01 | Paper Table 11. pass@4 94.3%, Coverage 93.4%, 42.6K tokens |
| GPT-6 Astra | OpenAI | 74.7% pass@1 | Codex, safeguards off, no cap, High, n=4 | 2026-10-01 | Paper Table 11. pass@4 92.4%, Coverage 93.2%, 38.7K tokens |
| GPT-6 Astra | OpenAI | 72.2% pass@1 | Codex, safeguards off, no cap, Medium, n=4 | 2026-10-01 | Paper Table 11. pass@4 90.1%, Coverage 90.3%, 34.2K tokens |
| GPT-6 Astra | OpenAI | 52.3% pass@1 | Codex, safeguards off, no cap, Low, n=4 | 2026-10-01 | Paper Table 11. pass@4 67.2%, Coverage 76.1%, 20.7K tokens |
| GPT-5.6 Cyber | OpenAI | 56.8% pass@1 | Codex, safeguards off, no cap, Max, n=4 | 2026-10-01 | Paper Table 11. pass@4 72.5%, Coverage 83.1%, 342.6K tokens |
| GPT-5.6 Cyber | OpenAI | 43.4% pass@1 | Codex, safeguards off, no cap, XHigh, n=4 | 2026-10-01 | Paper Table 11. pass@4 56.9%, Coverage 68.8%, 211.1K tokens |
| GPT-5.6 Cyber | OpenAI | 23.1% pass@1 | Codex, safeguards off, no cap, High, n=4 | 2026-10-01 | Paper Table 11. pass@4 34.0%, Coverage 43.9%, 127.7K tokens |
| GPT-5.6 Cyber | OpenAI | 10.5% pass@1 | Codex, safeguards off, no cap, Medium, n=4 | 2026-10-01 | Paper Table 11. pass@4 16.8%, Coverage 23.6%, 64.5K tokens |
| GPT-5.6 Sol | OpenAI | 55.8% pass@1 | Codex, safeguards off, no cap, Max, n=4 | 2026-10-01 | Paper Table 11. pass@4 68.7%, Coverage 83.9%, 241.5K tokens |
| GPT-5.6 Sol | OpenAI | 41.2% pass@1 | Codex, safeguards off, no cap, XHigh, n=4 | 2026-10-01 | Paper Table 11. pass@4 53.8%, Coverage 69.1%, 143.3K tokens |
| GPT-5.6 Sol | OpenAI | 16.7% pass@1 | Codex, safeguards off, no cap, High, n=4 | 2026-10-01 | Paper Table 11. pass@4 23.3%, Coverage 36.6%, 70.9K tokens |
| GPT-5.6 Sol | OpenAI | 5.8% pass@1 | Codex, safeguards off, no cap, Medium, n=4 | 2026-10-01 | Paper Table 11. pass@4 8.8%, Coverage 14.9%, 28.9K tokens |
Astra at Max leads at 88.0% pass@1, 31.2 points ahead of GPT-5.6 Cyber at Max and 32.2 ahead of Sol at Max. Astra at Medium already beats both of them at Max, at 72.2%, on roughly a seventh of Sol's output tokens per attempt. Sol swings the most with effort, from 5.8% at Medium to 55.8% at Max. For Sol, the effort setting is a bigger decision than the model name, a plain case of the test-time compute trade-off.
OpenAI published per-task costs for two of these models. Summed over every attempt at every effort level, Astra's 5,240 runs cost $83.4K and produced 215.7M output tokens, an average of $15.92 per attempt. Sol's 4,192 runs cost $161.9K and produced 506.9M output tokens, $38.62 per attempt. The paper's Anthropic and GPT-5.6 Cyber costs are estimates on a different basis, so this page does not compare cost across vendors.
Unconstrained Anthropic runs in Claude Code
Anthropic ran these internally in Claude Code at Max effort with safeguards off, no budget cap, four attempts per instance and a clarified malware prompt. They compare with each other, not with the Codex runs or the standardized harness.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 82.1% pass@1 | Claude Code, safeguards off, no cap, Max, n=4, clarified prompt | 2026-10-01 | Paper Table 12. pass@4 98.1%, Coverage@1 92.9%. 321.1K output tokens per attempt, from Table 11 |
| Claude Mythos 5.1 | Anthropic | 53.9% pass@1 | Claude Code, safeguards off, no cap, Max, n=4, clarified prompt | 2026-10-01 | Paper Table 12. pass@4 61.8%, Coverage@1 76.0%. No token figure published |
The effort ladder below is a separate single-attempt run of Opus 5.5. It compares across its own rows only, and its XHigh row is not a step toward the n=4 Max result above.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 77.9% pass@1 | Claude Code, safeguards off, XHigh, n=1, clarified prompt | 2026-10-01 | Paper Table 11. Coverage 89.3%, 240.1K output tokens per attempt |
| Claude Opus 5.5 | Anthropic | 66.8% pass@1 | Claude Code, safeguards off, High, n=1, clarified prompt | 2026-10-01 | Paper Table 11. Coverage 85.2%, 169.6K tokens |
| Claude Opus 5.5 | Anthropic | 59.9% pass@1 | Claude Code, safeguards off, Medium, n=1, clarified prompt | 2026-10-01 | Paper Table 11. Coverage 81.2%, 143.3K tokens |
| Claude Opus 5.5 | Anthropic | 53.1% pass@1 | Claude Code, safeguards off, Low, n=1, clarified prompt | 2026-10-01 | Paper Table 11. Coverage 84.4%, 117.5K tokens |
Opus 5.5 leads Mythos 5.1 by 28.2 points at Max. Put the Opus 5.5 rows next to its Vals entry and the gap between the two harnesses is stark. Safeguards off in Claude Code, it solves 82.1%. On the Vals board, where its own work accounts for 5.34% and fallback models carry the rest, it posts 33.59%. Astra's 88.0% and Opus 5.5's 82.1% come from different vendor harnesses, so they are not a head-to-head result.
Historical progression
SRE-Bench is new, and its only same-harness history is the generation steps inside the standardized table. Because these models ran in one harness with one set of caps, the steps are valid comparisons.
| Family | Earlier model | Solved of 262 | Later model | Solved of 262 | Change |
|---|---|---|---|---|---|
| OpenAI | GPT-5.5 | 3.82% | GPT-5.6 Sol | 30.53% | +26.7 points |
| OpenAI | GPT-5.6 Sol | 30.53% | GPT-6 Astra | 56.87% | +26.3 points |
| Anthropic | Claude Opus 5 | 12.21% | Claude Fable 5.1 | 22.90% | +10.7 points |
Source: paper Table 4, v2 dated 2026-10-01. Opus 5.5 is not in the standardized table, so the Anthropic line stops at Fable 5.1. GPT-6.1 Sol, Gemini 4 Argon, Opus 5.5 and Sonnet 5.5 appear only in Vals runs, and no earlier-version score exists for them on a matching setup.
OpenAI's two steps are about the same size, each adding roughly 26 points. Anthropic's single step is less than half that, and Fable 5.1's figure carries the fallback caveat above.
Failure modes and limitations
Hardened binaries
Protection presets are the obstacle that separates models. The table shows Score out of six on the standardized harness for unhardened and hardened builds, and the cost per point earned.
| Model | Unhardened Score | Hardened Score | Cost per point, unhardened | Cost per point, hardened | Cost multiple |
|---|---|---|---|---|---|
| GPT-6 Astra | 5.83 | 5.73 | $2.15 | $3.18 | 1.5x |
| GPT-5.6 Sol | 4.69 | 2.50 | $6.52 | $22.39 | 3.4x |
| Claude Fable 5.1 | 5.33 | 0.32 | $6.93 | $105.79 | 15.3x |
| Claude Opus 5 | Not printed | 0.33 | $7.22 | $66.98 | 9.3x |
| GPT-5.5 | Not printed | 0.11 | $11.02 | $138.54 | 12.6x |
| DeepSeek V4.1 Flash | Not printed | 0.08 | Not printed | Not printed | Not printed |
| Grok 4.5 | Not printed | 0.02 | $18.56 | $654.31 | 35.3x |
| GLM 5.2 | Not printed | 0.00 | Not printed | No points earned | Not applicable |
Source: paper Table 7 and section 4, v2 dated 2026-10-01. Astra's figures cover only its graded runs.
Fable 5.1 is the sharpest case. It scores 5.33 of 6 on unhardened builds, close to Astra, and 0.32 on hardened ones. Every model except Astra and Sol scores 0.33 or lower once protections go on. Astra drops 0.10 points. Sol keeps about half its Score, at a cost per point 3.4 times higher.
The unconstrained runs show the same pattern within each harness. In Codex, Astra loses 0 to 6 pass@1 points under any preset compared with its unhardened builds, while Sol loses 13.5 to 19.7. In Claude Code, Opus 5.5 loses 0 to 15 points. Mythos 5.1 loses under 4, but from an unhardened base of 53.9%.
Stripped symbols and optimization
Stripping symbols is the biggest build penalty on the standardized harness. The four weakest models lose 50 to 71% of their Score on stripped builds. Sol, Fable 5.1 and Opus 5 lose 10 to 15%. Astra loses nothing, 5.39 against 5.36. Turning on compiler optimization costs the weaker models 17 to 41%.
Games, especially in Rust
In the unconstrained runs, games are the hardest domain for GPT-5.6 Sol in Codex and for Mythos 5.1 in Claude Code. Sol solves 1.6% of Rust game instances in Codex, and Mythos 5.1 solves none in Claude Code. Rust is the lowest game cell for all four models. Astra ranges from 35.9% on Rust to 90.6% on C, and Opus 5.5 from 34.4% on Rust to 76.6% on C.
Source-level security skill does not transfer
GPT-5.6 Sol and Fable 5.1 both have strong source-level security results, and on the standardized harness they fully solve 31.5% and 26.9% of their graded instances. Reading code and reading a stripped, protected binary are different skills, and a source-level security score does not predict this one.
Refusals
Safeguard refusals left only 181 of Astra's 262 standardized instances, 69.1%, graded, even though the tasks do not involve exploitation. The paper says this raises concerns about Astra's practical use for reverse engineering. The Claude models hit the same wall from a different direction. Their Vals scores depend on fallback to earlier models, with Opus 5.5 falling from 33.59% to 5.34%, Fable 5.1 from 22.90% to 10.69% and Sonnet 5.5 from 30.15% to 19.08% when fallback tasks count as failures.
Picking the right attempt
Astra's 99.2% pass@4 at Max looks close to a solved benchmark. It isn't. Astra solved 191 instances in all four attempts, 33 in exactly three, 23 in exactly two, 13 in exactly one and 2 in none. That leaves 69 instances where right and wrong submissions sit side by side, and 126 of 1,048 attempts fail. SRE-Bench does not test whether an agent can pick the correct submission without the private grader, and in production there is no private grader. Plan on the 88.0% pass@1.
Prompt wording
The malware prompt's wording moves Opus 5.5 at Max in Claude Code from 76.3% pass@1 with the original prompt to 82.1% with a clarified one, and Coverage@1 from 87.4% to 92.9%. Mythos 5.1 barely moves, 53.3% to 53.9%. The paper publishes neither prompt, so an outside team cannot reproduce the difference.
Reward hacking
An early build leaked a hidden behavior's name as a readable string, and an agent found the trigger by scanning strings instead of analyzing the code. The authors removed the string and added a regression check that fails the build if it happens again. They used Opus 4.8 and GPT-6 Astra for audits.
Breadth and access
Nineteen programs averaging 16.9K lines is far larger than earlier RE sets but far smaller than a million-line system. The binaries are private, so outside parties cannot re-run the benchmark locally or check a vendor's number themselves.
How SRE-Bench compares to related benchmarks
ProgramBench also starts from a compiled binary, but it uses known open-source projects and asks the agent to reimplement the program's behavior. SRE-Bench uses private binaries and grades the recovered semantics with task-specific graders for protocol coverage, byte-exact decoding, malware cleanup and hidden triggers. ProgramBench tests whether a model can rebuild a program. SRE-Bench tests whether it understands one it cannot read.
Against the CTF-derived sets, NYU CTF Bench, Cybench, CTF-Dojo and CTFTiny, and the toy RE sets, CREBench, AgentRE-Bench and CrackMeBench, the difference is scale and provenance. Those use public challenges or programs of 36 to 526 average lines with easy protections. SRE-Bench averages 16,915.8 lines, stays private and adds hardening. The contamination test above shows what that buys. Models that ace a small or public-lineage compressor drop sharply on the real-scale private one.
ExploitBench, ExploitGym and SEC-Bench Pro cover exploit development and other security work. SRE-Bench leaves exploit development out on purpose, so a score measures reverse engineering without mixing in exploitation skill. Teams assessing offensive security capability, or planning red-teaming work, need both kinds of result.
SWE-bench tests source-level fixes in public repositories. SRE-Bench is the binary-first counterpart, with private programs and no source code to read. A team that picked a model on SWE-bench has learned nothing yet about how it handles a stripped, protected binary.
What it means for teams choosing a model
GPT-6 Astra is the pick on the standardized harness. It solves 56.87% of all 262 instances, against 30.53% for GPT-5.6 Sol and 22.90% for Claude Fable 5.1. It does it for $13.50 per graded instance, against $42.50 for Sol and $34.70 for Fable 5.1, in 18 minutes, the fastest in the table. Only DeepSeek V4.1 Flash is cheaper, and it solves 0.76%.
Budget for refusals. Astra was graded on only 181 of 262 instances because safeguards refused the rest. The 56.87% already counts those as misses, but a team doing reverse engineering work needs a fallback path or a vendor access program for the instances the model will not touch.
Test on hardened binaries if your binaries are hardened. Astra holds at 5.73 of 6 on hardened builds. Sol drops to 2.50, Opus 5 to 0.33, Fable 5.1 to 0.32 from 5.33 unhardened, and GPT-5.5 to 0.11. Astra's hardened Score covers only graded runs, so rank on Solved of 262 to compare. No hardened data exists for GPT-6.1 Sol, Gemini 4 Argon or Opus 5.5.
Read Vals scores with the fallback share next to them. On the Vals board, GPT-6.1 Sol posts 50.76% at $2.69 per test, Gemini 4 Argon 44.27% at $30.45 and Claude Opus 5.5 33.59% at $34.65. Each ran on its own date with its own settings. Opus 5.5's own work accounts for 5.34%, and the rest passes through fallback to older Claude models. A buyer comparing Vals scores needs to know how much of each one the named model earned.
Treat effort as part of the model. In Codex, Astra at Medium reaches 72.2% pass@1 on 34.2K output tokens per attempt, while Sol at Max needs 241.5K tokens for 55.8%. In Claude Code, Opus 5.5 climbs from 53.1% at Low on 117.5K tokens to 77.9% at XHigh on 240.1K, in a single-attempt run. Output tokens per attempt differ by up to seven times across these settings, so weigh cost per token and token use together at the effort you plan to run.
Plan on pass@1. Astra's 99.2% pass@4 needs a selector that SRE-Bench does not include, and 69 of its instances mix right and wrong attempts. The number to plan around is 88.0%, and that figure comes from an unconstrained harness customers cannot buy.
Cheap models do not substitute. DeepSeek V4.1 Flash costs $0.50 per instance and solves 0.76%. GLM 5.2 costs $26.60 and solves none. Only the top three frontier models in the standardized table solve more than an eighth of the set.
You can compare these models across other benchmarks on the Klu LLM leaderboard. To run the same kind of deterministic, task-graded evaluation on your own engineering workflows, see how teams set it up in Klu on the engineering page.