SEC-Bench Pro

A vulnerability-discovery benchmark that asks a coding agent to find a disclosed bug in V8, SpiderMonkey, or the Linux kernel and prove it with a crashing proof-of-concept

Security and operationsSolved instances, three-image judge

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 14 min read

What is SEC-Bench Pro?

SEC-Bench Pro is a vulnerability-discovery benchmark for coding agents, including those built on frontier models. Each task points the agent at a real, already-disclosed security bug somewhere inside a large production codebase, and the agent has to prove the bug exists by writing a proof-of-concept (PoC) input that crashes the instrumented build in the right place for the right reason. The targets are Google's V8 JavaScript engine, Mozilla's SpiderMonkey and the Linux kernel.

The University of Illinois Urbana-Champaign built it. Hwiwon Lee is first author, with Chunqiu Steven Xia and Lingming Zhang as senior authors. The paper is arXiv 2605.26548, and the maintainers run a public leaderboard at sec-bench.github.io. It follows the same group's earlier SEC-bench.

The benchmark scores discovery plus a working crash. It does not score exploitation. An agent that crashes V8 with a use-after-free gets full credit without ever turning that crash into code execution. Exploit development is what ExploitBench and ExploitGym measure, and the difference shapes how you read every number below.

The authors built it because they think older vulnerability benchmarks such as ARVO and CyberGym test the wrong skill. Those benchmarks hand the agent a fuzz harness with a narrow entry point, so the agent mostly mutates inputs instead of auditing code. They grade by matching sanitizer crashes, which credits crashes in the wrong place. And their targets are small. SEC-Bench Pro drops the fuzz harness, uses three large production codebases, and checks every PoC against three builds.

How the tasks work

The current dataset (paper v2) has 344 validated vulnerabilities.

TargetInstancesWhere they come fromPoC format
V810386 are bounty-qualified, with $1,540,750 in cumulative Google VRP awardsJavaScript
SpiderMonkey104Mozilla Bugzilla sec-high and sec-critical reports (the leaderboard tab says "Firefox")JavaScript
Linux kernel137CVE-backed, 89 from kernelCTF and 48 from syzbotC program run in a KASAN kernel under QEMU

The browser-engine bug classes are type confusion, use-after-free, sandbox bypass, out-of-bounds read and write, integer overflow or truncation, and JIT or code-generation bugs. The kernel set adds subsystem bugs and race conditions.

The agent gets the vulnerable source tree, the relevant source paths, an instrumented binary, the permitted runtime flags, the vulnerability class and the expected error type. It does not get the reference PoC, the root-cause location, the patch, the crash trace, the original report or the fixed source. The prompt tells it to work autonomously and forbids running fuzzers. So the agent knows roughly where to look and what kind of bug to expect, and it has to read the code, work out the root cause and build an input that reaches it.

On the kernel, privilege matters. 98 of the 137 Linux instances are reachable by an unprivileged user and are graded at uid 1000. The other 39 need a capability in the init namespace and are graded as root.

Harness

Each instance ships three Docker images. The vulnerable image is the reported revision built with sanitizers, or KASAN for the kernel. The fixed image adds the linked patch. The latest image has every later upstream fix applied.

Every execution runs inside a Docker sandbox. The maintainers run each vendor's agent unmodified with its default tools: Codex for OpenAI models, Claude Code for Anthropic models and OpenCode for open-weight models. That means the leaderboard measures agent-plus-model pairs. "GPT-5.5" on this page always means GPT-5.5 inside Codex.

Each instance gets a 90-minute wall-clock budget. Each PoC execution gets 300 seconds, with a fixed extra buffer on Linux for the QEMU boot, and up to three retries against the same image. Provider-side web search is off in the final harness, and the maintainers check trajectories for completed shell network commands. That offline rule came from a real leak, covered under failure modes.

Scoring

An LLM judge looks at each candidate PoC's behavior on all three images and labels it Verified, Unsure or Illegal. Verified means the PoC crashes the vulnerable image, the crash matches the target source and vulnerability type, and the fixed and latest images do not contradict that attribution. An instance is solved when at least one candidate PoC is Verified. The default configuration also counts Unsure as a success. All 22 Unsure outcomes in the paper are Linux candidates, and manual review resolved 20 of them as verified and 2 as illegal.

The headline number is solved instances divided by all instances. A run that hits the 90-minute cap counts as a failure, which turns out to decide the Anthropic result.

Against manually adjudicated ground truth, the judge credited 455 of 464 solved run-instances, with 4 false positives and 13 false negatives. That is 99.1% precision and 97.2% recall. The three-image check is what makes the benchmark stricter than crash matching, and the paper's Table III shows by how much on V8:

ConfigurationCrash on vulnerable imagePlus clean fixed imageThree-image judge
Codex GPT-5.5581849
Codex GPT-5.4421636
Claude Code Opus 4.6421723
SEC-Bench Pro maintainersV8V8 instances credited by each grader

A crash-only grader would put Opus 4.6 level with GPT-5.4 on V8. The judge credits only 23 of Opus 4.6's 42 crashes, against 49 of GPT-5.5's 58. Crash-only grading flatters the weaker agent most. Requiring a fully clean fixed image undercounts in the other direction, crediting GPT-5.5 with only 18.

Versions

VersionPaperInstancesTargetsWhat changed
260505 leaderboard snapshotv1, 2026-05-26183V8 (103), SpiderMonkey (80)V8 leaderboard opened 2026-05-01, SpiderMonkey on 2026-05-05
260617 leaderboard snapshotv2, 2026-07-20344V8, SpiderMonkey (104), Linux (137)Adds Linux, privilege-aware kernel grading and the offline harness; overall and Linux leaderboards opened 2026-06-17

OpenAI's system cards still use "the May 2026 version", the 183-instance V8 and SpiderMonkey set with no Linux tasks, and a grader OpenAI modified itself. That is why the leaderboard below splits into two tables.

Current leaderboard

Maintainers' run, 344 instances

These results come from the paper v2 Table II, and the Overall tab of sec-bench.github.io at version 260617 prints the same counts. Every row uses the three-image judge, the 90-minute budget and the offline harness. The maintainers reached every model except OpenAI's through AWS Bedrock.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-5.5OpenAI58.4% (201/344)Codex, xhigh2026-07-20 (snapshot 260617)arXiv v2 Table II; 14 runs timed out
GPT-5.4OpenAI39.0% (134/344)Codex, xhigh2026-07-20 (snapshot 260617)arXiv v2 Table II; 18 runs timed out
Opus 4.6Anthropic30.8% (106/344)Claude Code, max2026-07-20 (snapshot 260617)arXiv v2 Table II; 208 runs timed out
GLM-5Z.ai3.8% (13/344)OpenCode, high2026-07-20 (snapshot 260617)arXiv v2 Table II; 30 runs timed out
Kimi K2.5Moonshot AI2.3% (8/344)OpenCode, high2026-07-20 (snapshot 260617)arXiv v2 Table II; 3 runs timed out
MiniMax M2.5MiniMax0.6% (2/344)OpenCode, high2026-07-20 (snapshot 260617)arXiv v2 Table II; 14 runs timed out
SEC-Bench Pro maintainersCodex, Claude Code, and OpenCode344 instances, snapshot 260617Solved instances, three-image judgearxiv.org

Comparable with every row in this table, since they share one grader, budget and instance set. Not comparable to the OpenAI system-card table below or to the 183-instance v1 results.

The per-target split shows where each configuration earns its score:

ModelV8 (103)SpiderMonkey (104)Linux (137)
GPT-5.547.6% (49)44.2% (46)77.4% (106)
GPT-5.435.0% (36)25.0% (26)52.6% (72)
Opus 4.622.3% (23)27.9% (29)39.4% (54)
GLM-51.9% (2)5.8% (6)3.6% (5)
Kimi K2.51.9% (2)2.9% (3)2.2% (3)
MiniMax M2.50.0% (0)0.0% (0)1.5% (2)
SEC-Bench Pro maintainersCodex, Claude Code, and OpenCode344 instances, snapshot 260617Solved instances by target, three-image judgearxiv.org

GPT-5.5 in Codex leads by 19.4 points over GPT-5.4 and by 27.6 points over Opus 4.6 in Claude Code, and it leads on every target. The open-weight models in OpenCode barely register.

The Opus 4.6 number needs one more fact. Claude Code completed only 136 of its 344 runs inside the budget (36 on V8, 43 on SpiderMonkey, 57 on Linux). On completed runs alone, Opus 4.6 solves 61.1% of V8, 58.1% of SpiderMonkey and 91.2% of Linux. GPT-5.5's completed-only rates are 48.0%, 44.7% and 83.2%. Opus 4.6 is accurate when it finishes and usually does not finish. The paper keeps the all-instances rate as the headline because every configuration gets the same 90 minutes and, in the authors' words, "finishing within that budget is part of the task." I agree with that call. A security team running an agent overnight still has a clock.

There is no score for any Anthropic or Google model newer than Opus 4.6. The leaderboard footnote says newer Claude Opus and Mythos results are unavailable "due to safeguard restrictions", and the paper explains that Bedrock access sits outside the verification program that authorizes newer Claude models for offensive-security work, so guardrails block the runs. None of the Opus 5.5, Fable 5.1, Mythos 5.1 or Gemini 4 documents report a SEC-Bench Pro result.

SEC-Bench Pro: solved vs time per taskMaintainers' run, 344 instances, three-image judge, snapshot 260617
20%40%60%20 min50 min1.67 hr
FrontierOtherOpenAIBehind the frontier
Mean time per task, log scale
Source: SEC-Bench Pro paper, arXiv 2605.26548 v2 (2026-07-20), Table II, accessed 2026-10-06. Score is solved instances over all 344. Time is the instance-weighted mean of the per-target Avg. Runtime column, (103 x V8 + 104 x SpiderMonkey + 137 x Linux) / 344. Timed-out runs count at the 90-minute cap, so Opus 4.6 sits near it. The paper does not publish a numeric cost per task for this table.

Read the chart as a time frontier. GPT-5.5 is the highest point and the second fastest, at 24.4 minutes per task. Kimi K2.5 is faster at 21.2 minutes and solves 8 instances. Opus 4.6 takes 68.6 minutes per task for about half of GPT-5.5's score. GPT-5.4 is slower than GPT-5.5 and solves fewer, so the upgrade improved both axes at once.

The paper publishes cost per solved instance for three configurations, priced at provider list rates:

ConfigurationCost per solved instanceNote
Codex GPT-5.5$19.23Cheapest per success
Claude Code Opus 4.6$71.583.7x GPT-5.5; 2.5x to 4.5x Codex per success by target
OpenCode GLM-5over $450 (V8 only)High because it solves only 2 V8 instances
SEC-Bench Pro maintainersCost per solved instance

The paper's own section heading for this is "Higher cost does not buy capability", and the numbers back it.

OpenAI system cards, 183 instances

OpenAI reports SEC-Bench Pro in its system cards on the May 2026 version. For the GPT-6 Astra card, OpenAI found the public grader "insufficient to accurately root-cause target paths, thus leading to an error ceiling." It added a second agent that audits the solver's trajectory and checks whether repairing the target source files invalidates the PoC, and it reported the change to the maintainers. OpenAI plots pass@1 against mean total solution length in tokens, which makes each chart a test-time compute curve.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI85.4% pass@1Scaffold unpublished, root-cause grader; effort not stated2026-09-03GPT-6 Astra System Card, Appendix A.8.1.2, PDF p.148
GPT-5.6 SolOpenAI79.1% pass@1Same; maximum point on its curve, using more tokens2026-09-03Astra card, Appendix A.8.1.2
GPT-6.1 SolOpenAI78.8% pass@1Same2026-09-29GPT-6.1 Sol addendum, PDF p.44, Fig. 36
GPT-6 SolOpenAI66.3% pass@1Same2026-09-03Astra card, Appendix A.8.1.2
GPT-6 LunaOpenAI34.2% pass@1Same2026-09-03Astra card, Appendix A.8.1.2
OpenAI183 instances, May 2026 versionpass@1, OpenAI root-cause grader

Comparable with other rows in this table only. It has no Linux tasks, a vendor-modified grader, an unpublished scaffold and vendor-run pass@1, so none of these numbers line up with the 344-instance table.

GPT-6 Astra leads at 85.4%, 6.3 points ahead of GPT-5.6 Sol, 6.6 ahead of GPT-6.1 Sol and 19.1 ahead of GPT-6 Sol. The more useful signal is in the card text about tokens. The Astra card says Astra is "significantly more token efficient and more capable at vulnerability identification and exploit development" than GPT-5.6 Sol. It says GPT-6 Sol "performs broadly comparably to GPT-5.6 Sol at similar output-token counts, although GPT-5.6 Sol reaches a higher maximum score of 79.1% using more tokens." The addendum says GPT-6.1 Sol "achieves a similar peak score to GPT-5.6 Sol at a substantially shorter mean solution length, while remaining below GPT-6 Astra."

OpenAI has not published the step limit, wall-clock cap, scaffold, tool list, the reasoning effort behind each plotted point, the number of samples behind pass@1, or any dollar cost per task. The earlier GPT-5.6 card (2026-07-09) used the public grader and printed its SEC-Bench Pro curves with no numeric labels, so it gives no numbers to compare.

Historical results

The v1 paper (2026-05-26) ran three configurations on the 183-instance set with the same three-image judge and reported mean cost per instance.

ModelOrganizationTargetVerifiedHarness and setupMean cost per instanceFrontier
GPT-5.4OpenAIV8 (103)33 (32.0%)Codex$9.97Frontier
Opus 4.6AnthropicV8 (103)22 (21.4%)Claude Code$17.93
Kimi-K2.6Moonshot AIV8 (103)12 (11.7%)OpenCode$6.88Frontier
Opus 4.6AnthropicSpiderMonkey (80)31 (38.8%)Claude Code$15.22Frontier
GPT-5.4OpenAISpiderMonkey (80)19 (23.8%)Codex$8.20Frontier
SEC-Bench Pro maintainersCodex, Claude Code, and OpenCode183 instances (v1)Verified instances, three-image judgearxiv.org

Comparable with rows for the same target here. Not comparable to the 344-instance table, because the SpiderMonkey set grew from 80 to 104 instances, Linux did not exist yet, and v1 used Kimi-K2.6 where v2 uses K2.5. Source: arXiv v1 Tables 3 and 4.

In v1, Codex won V8 by 11 instances and Claude Code won SpiderMonkey by 12. Together they covered 39 of 103 V8 bugs and 39 of 80 SpiderMonkey bugs, agreeing on only 16 and 11. In the v2 run, Opus 4.6 still edges GPT-5.4 on SpiderMonkey (29 to 26), but GPT-5.5 solves 46. Within the v2 run, that upgrade took Codex from 39.0% to 58.4% overall: V8 from 36 to 49 solved, SpiderMonkey from 26 to 46, Linux from 72 to 106. Linux alone added 34 instances, and the two browser engines added 33 together.

Failure modes

Most failures never crash anything. The paper sorts 3,486 illegal verdicts. On the browser engines, "No crash" and "Generic JS" errors together make up over 95% of them. Only 132 of the 3,486 produce a real crash on the vulnerable image, and 119 of those still fail, because the crash lands outside the target source, has the wrong sanitizer class, or also happens on the fixed image. Agents mostly fail at reaching the bug, not at aiming once they get there.

Open-weight models stall before any crash signal. Only 9 of the open-weight models' 2,179 illegal PoCs crash at all, off-target or in the wrong class, against 85 of 1,298 for the closed-weight configurations. Open-weight models use the full 90 minutes on 4.6% of instances, against 23.3% for closed-weight ones.

Timeouts decide the Claude Code rank. 208 of Opus 4.6's 344 runs hit the cap. Claude Code's failed runs average over 80 minutes and its successes under 40. Codex fails differently: GPT-5.5's failed runs use over twice the tokens of its successes.

Two search styles. Codex with GPT-5.5 solves 48 of its 49 verified V8 instances with a single PoC. Claude Code submits 927 candidate PoCs on V8 against 100 for Codex, and 99.1% of its V8 illegal verdicts never crash the engine. That spray still finds things. Claude Code solves 9 SpiderMonkey instances GPT-5.5 misses, the two agree on only 20, and their union reaches 55 against 46 for GPT-5.5 alone.

The kernel is easier than the browsers. GPT-5.5 reaches 77.4% on Linux, 47.6% on V8 and 44.2% on SpiderMonkey. The paper puts this down to reachability. A kernel bug is reached through a structured syscall interface, while a browser bug needs a JavaScript program that pushes the engine through specific JIT and garbage-collection states.

Online leakage. In earlier runs with network access and root grading, Codex solved 122 of 137 Linux instances with GPT-5.5 and 105 with GPT-5.4 by fetching evidence the task withholds. For CVE-2025-38001 it found the exact upstream fix commit, for CVE-2025-38177 the exact patch, and GPT-5.4 retrieved the original syzkaller reproducer for CVE-2022-50367. The final harness disables web search and reruns Linux offline. The paper says those earlier scores "do not form a matched network ablation", so they are not an online-versus-offline comparison either.

Privilege shortcuts. Replaying PoCs from a root-only evaluation at uid 1000, 49 of GPT-5.5's 73 successes on user-reachable Linux instances and 43 of GPT-5.4's 46 stop reproducing. In the final privilege-aware runs, GPT-5.5 solves 78 of the 98 user-reachable instances and GPT-5.4 solves 48. Of those successful PoCs, 68 of GPT-5.5's 78 and 43 of GPT-5.4's 48 call unshare(CLONE_NEWUSER) to create a user namespace. One technique accounts for most of the user-level kernel successes.

Grader dependence. In v1, crash-only grading inflates verified counts by 1.21x to 1.73x: Codex on V8 1.36x, Claude Code on V8 1.73x, Claude Code on SpiderMonkey 1.48x, Codex on SpiderMonkey 1.21x. OpenAI says the public grader has an error ceiling and replaced it. The grader is part of the score.

Contamination. Every task is a disclosed bug with a public fix, so it can sit in a model's training data. The maintainers' answer is a self-evolving pipeline that adds new tasks as new reports are disclosed. Neither the paper nor OpenAI's cards report a held-out, post-cutoff SEC-Bench Pro split, and the cards' contamination checks cover other benchmarks. Nobody has published a contamination audit of SEC-Bench Pro.

Coverage and judge. The authors name the representativeness of 344 bugs across three targets as the main external-validity threat and the LLM judge as the main internal one. On difficulty by class, v1 found integer overflow and truncation easiest on V8 (Codex and Claude Code both 4 of 6) and JIT or code-generation bugs unsolved (0 of 2 for all three V8 agents). Codex and Claude Code rank classes almost identically (Spearman 0.88).

Real discoveries. The evaluation runs turned up two zero-day vulnerabilities in V8 and a duplicate bug in SpiderMonkey. One V8 finding, a BigInt heap overflow sandbox bypass, was taken to arbitrary code execution and earned a $20,000 Chrome VRP bounty.

BenchmarkWhat the agent getsWhat counts as success
SEC-Bench ProSource tree, target paths, bug class, instrumented binary; no patch or reportA crashing PoC attributed to the target bug across three builds
ExploitBench41 V8 N-days with source, git history through the fix, bug description, patch diff, both binaries"Cap Percent" over 16 capability flags in five tiers up to arbitrary code execution, five attempts
ExploitGym869 challenges (502 userspace C/C++, 181 V8, 186 kernel), each with a PoV that already triggers the bugRetrieve a flag via remote code execution, with a judge confirming the intended bug
SRE-Bench262 binaries from 19 privately developed programs, no sourceAll six reverse-engineering objectives met
ARVO and CyberGymOSS-Fuzz reports and fuzz harnessesSanitizer crash match
SWE-benchRepository and issue textA patch that passes the tests

The security benchmarks form a pipeline. SEC-Bench Pro asks the agent to find the trigger without the patch. ExploitBench gives it the patch and asks how far toward code execution it gets. ExploitGym hands it the trigger and asks for a working exploit. On ExploitGym, OpenAI reports GPT-6 Astra at 42.4% per attempt, GPT-6.1 Sol at 35.1%, GPT-5.6 Sol at 30.3% and GPT-6 Sol at 22.1%. SRE-Bench removes the source code and uses private programs to avoid contamination, which is the weakness SEC-Bench Pro has to manage by adding new bugs.

SWE-bench is the closest coding relative. Both start from a real bug in a real repository, but SWE-bench asks for a fix and SEC-Bench Pro asks for an input that breaks the code. The v1 paper compares its ranking to SWE-Bench Pro, "where the strongest frontier model stays below 45% (Pass@1)" on long-horizon bug fixing. For more on how these suites fit together, see LLM benchmarks and LLM evaluation.

What it means for teams choosing a model

If you want an agent to audit large C++ or kernel code for real bugs, Codex with GPT-5.5 at xhigh is the best independently measured option. It solves 58.4% of the 344 instances against 30.8% for Claude Code with Opus 4.6, and it leads on V8, SpiderMonkey and Linux. It is also the fastest of the closed-weight configurations at 24.4 minutes per task and the cheapest per success at $19.23, against $71.58 for Opus 4.6.

Claude Code with Opus 4.6 is held back by time, not accuracy. It completes only 136 of 344 runs within 90 minutes but solves most of what it completes. A team with a tighter budget than 90 minutes should expect a lower Claude Code score than the table shows. Newer Anthropic models have no score at all, since safeguards block the maintainers' Bedrock runs.

If coverage matters more than cost, run both agents. On SpiderMonkey, GPT-5.5 and Opus 4.6 agree on only 20 solved bugs and together solve 55, against 46 for GPT-5.5 alone. On V8 the union adds only 2 instances (51 against 49), and on Linux GPT-5.5 covers almost everything Opus 4.6 solves. Pairing pays off on SpiderMonkey and barely anywhere else.

OpenAI's own numbers put GPT-6 Astra at 85.4% and GPT-6.1 Sol at 78.8% on the smaller set. Compare those with each other, not with the maintainers' table, because the grader and the instance set differ. The cards support a statement about token length, that GPT-6.1 Sol reaches GPT-5.6 Sol's peak with shorter solutions, and say nothing about price. OpenAI treats GPT-6 Astra and GPT-6.1 Sol as Critical in cybersecurity under its Preparedness Framework, so access to those models is gated.

Skip the open-weight models in OpenCode for this kind of work. GLM-5, Kimi K2.5 and MiniMax M2.5 solve 3.8% or less. Their PoCs rarely produce a crash at all.

Last, treat any vulnerability-finding score as tied to its grader. A crash count overstates the weaker agent the most, and a vendor number only compares with another number from the same card. If you are measuring this yourself, borrow the three-image idea: test each finding against a patched build before you count it. That check belongs in any red-teaming workflow that uses agents to find bugs.

To compare these models on your own code, start from the LLM leaderboard and run the same find-and-verify eval on your own repositories with Klu's engineering workflows.