DeepSWE

Datacurve's coding-agent benchmark of 113 original, long-horizon engineering tasks across 91 open-source repositories, graded by hand-written functional verifiers

Agentic codingpass@1

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 15 min read

What is DeepSWE v1.1?

DeepSWE is a benchmark for coding agents built by Datacurve. It has 113 software engineering tasks spread across 91 actively maintained open-source repositories in TypeScript, Go, Python, JavaScript and Rust. Wenqi Huang, Charley Lee, Leonard Tng and Serena Ge describe it in the paper DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks. Version 1 shipped in May 2026. Version 1.1, released on 2026-06-14, keeps the same 113 tasks and changes how they are graded. The official leaderboard lives at deepswe.datacurve.ai.

Datacurve built it to fix two problems with SWE-bench-style evaluation. The first is contamination. SWE-bench tasks come from merged pull requests, so the fix sits in public commit history and in pretraining corpora. The second is grading. An inherited test suite was written to check one specific human fix, so it can fail a different solution that is also correct. DeepSWE's answer to both is to write every task from scratch, never merge it upstream, and grade it with a verifier written by hand to check observable behavior.

The design produces wider separation between models. On the eight models with public reports on both benchmarks, the paper measures a 69.8-point spread on DeepSWE against 29.7 points on SWE-bench Pro. Its verifier audit found 1.4% disagreement with an LLM judge on DeepSWE, against 32.4% on SWE-bench Pro.

On Datacurve's board snapshot of 2026-09-22, GPT-6 Astra at xhigh effort leads with 74.1% pass@1. Gemini 3.8 Flash at high effort scores 73.8% for about half the cost per task, and Claude Opus 5 at max effort scores 73.6%. The run-to-run confidence intervals of the top five rows overlap, so the order among them is not established. Vendor-reported numbers go higher. Google claims 77.9% for Gemini 4 Argon, measured outside Datacurve's harness.

How the tasks are built

Each task is a short prompt, an executable verifier and a reference solution. Prompts run about 2,000 characters, roughly half the length of a SWE-bench Pro prompt. The reference solutions are 5.5 times longer in lines than SWE-bench Pro's and touch multiple files. The agent gets less instruction for a bigger change than SWE-bench Pro asks for.

Repositories have to be public, actively maintained, permissively licensed and above 500 GitHub stars. The median repository contributes one task, so no single codebase dominates the score. The PROVENANCE.md file in the deep-swe repo lists the upstream project and license behind every task, including helm/helm, fastapi/fastapi, httpx, KaTeX, go-git, bandit and drizzle-orm. Datacurve's Apache-2.0 license covers only its own specs, harness and verifiers.

Authoring has a few quality gates worth knowing about. Every verifier runs three times during authoring, and any task with a flaky test goes back for revision. Reviewers check that the prompt and verifier map one to one, so the verifier tests nothing the prompt does not ask for and the prompt asks for nothing the verifier skips. They check acceptance breadth, meaning the verifier accepts any correct implementation, along with realism and a clean environment. Ambiguous requests are excluded, so every task has one defensible solution.

The harness and environment

The official board runs every model in mini-swe-agent, a model-agnostic agent loop with a single bash tool, pinned to commit adfe2023. Using one minimal harness for every model removes scaffold quality from the comparison. It also means the board does not measure Claude Code, Codex CLI or any other product harness. A 10-task SWE-bench Pro pilot in the paper found no systematic handicap from mini-swe-agent. Opus 4.7 scored 50% in it against 40% in its native harness, GPT-5.5 scored 40% in both, and Gemini 3.1 Pro scored 40% against 20%. Ten tasks per model is a small pilot.

Agents run in sandboxed containers with no internet access, starting from a shallow git clone at the base commit. The shallow clone closes a hole that has hurt SWE-bench Pro, where agents recover the merged fix by reading .git history. Wall-clock timeout is 9,000 seconds, or 2.5 hours. Only 67 of 7,174 scored rollouts hit it, 0.9%. The paper states no step limit, and the context limit is each model's native context window.

Mean time per attempt on the configurations in the leaderboard below runs from about 10 to 80 minutes. These are long tasks by the standard of most coding benchmarks, but short next to multi-hour suites such as FrontierSWE v2.

How scoring works

A rollout passes only if every verifier test passes. There is no partial credit. Datacurve runs four rollouts per task, about 452 scored attempts per configuration, where a configuration is one harness, one model and one reasoning effort.

  • pass@1 is the attempt pass rate, macro-averaged per task so every task counts equally regardless of how many attempts scored.
  • pass@4 is the share of tasks with at least one passing rollout out of four.
  • 95% CI is run-to-run: 1.96 times the standard deviation across runs, divided by the square root of the number of runs.

What counts as a failure matters for reading the board. Provider, verifier and network errors are excluded from the denominator. Context-window exhaustion and agent timeouts count as failures. A model that burns through its context on a long task loses the attempt.

Every row also carries mean cost, tokens, steps and wall-clock time. Datacurve computes cost from measured token counts at each provider's list price and prints the price basis per row. For GPT-6 Astra that is $10 per million uncached input tokens, $12.50 per million cache writes, $1 per million cache reads and $50 per million output tokens. On Astra at xhigh, output tokens account for $1.48 of the $4.43 per-attempt cost and cached input for the rest.

Datacurve publishes the tasks, verifiers, trajectories, patches, verifier outputs and a trajectory browser. It has not released the judge prompt used in its verifier audit.

Versions

VersionReleasedScoringBoard
v1May 2026Exit code of the test runFrozen, 29 configurations, generated 2026-06-20
v1.12026-06-14Patch graded in a separate clean container, structured CTRF reports scored per test node IDLive, 70 configurations, snapshot 2026-09-22

The v1.1 release post lists three changes. The agent now commits its work, and Datacurve extracts the git patch and grades it in a fresh container, so nothing the agent did to the environment outside the patch can help it pass. Test results come back as structured CTRF reports scored by test node ID instead of a single exit code. Agents work on a normal main branch with future commits hidden, instead of a detached HEAD. Datacurve also fixed dependency drift and removed flaky tests on some tasks.

The re-grade barely moved the aggregate. Across the nine configurations run on both versions, the pooled pass rate went from 53.46% on v1 to 53.40% on v1.1, according to Datacurve's delta file. Individual configurations moved more, from -3.8 points for GPT-5.4 xhigh to +6.0 for GPT-5.5 medium. Individual tasks swung hard: vulture-persistent-analysis-cache fell from 83% to 17%, and narwhals-rolling-window-suite rose from 33% to 100%. A v1 score and a v1.1 score for the same model are different measurements.

Current leaderboard

Three groups of results exist for v1.1, and they do not mix. Datacurve's own board is the only one where every row shares a harness, runner and trial count. Mercor ran the same tasks with its own settings. Google and Anthropic reported their own numbers.

Datacurve official board, v1.1, mini-swe-agent

Datacurve ran every row in mini-swe-agent with four rollouts per task; rows are comparable with each other and not with Mercor's run or vendor-reported numbers. Source for every row is the v1.1 leaderboard data, board snapshot 2026-09-22, latest job finished 2026-09-01. The table shows 31 of the board's 70 configurations: the top configurations by pass@1, the best row of each lower-ranked model line, and GPT-5.4 xhigh for the progression section. Cost and time are means per attempt. Claude Fable 5 rows have 430 to 436 scored attempts against about 452 for most configurations.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI74.1%mini-swe-agent, xhigh effort2026-09-22pass@4 80.5%, CI ±2.9, $4.43, 18.9 min, 29 steps
Gemini 3.8 FlashGoogle73.8%mini-swe-agent, high effort2026-09-22pass@4 85.8%, CI ±1.4, $2.36, 11.4 min, 166 steps
Claude Opus 5Anthropic73.6%mini-swe-agent, max effort2026-09-22pass@4 88.5%, CI ±3.9, $11.84, 31.9 min, 99 steps
GPT-6 AstraOpenAI73.2%mini-swe-agent, high effort2026-09-22pass@4 82.3%, CI ±3.4, $3.92, 17.3 min, 27 steps
GPT-6 AstraOpenAI73.2%mini-swe-agent, max effort2026-09-22pass@4 79.6%, CI ±0.8, $7.50, 33.0 min, 28 steps
Claude Opus 5Anthropic73.2%mini-swe-agent, xhigh effort2026-09-22pass@4 85.8%, CI ±3.1, $9.07, 26.0 min, 89 steps
Claude Opus 5Anthropic72.8%mini-swe-agent, high effort2026-09-22pass@4 87.6%, CI ±1.9, $6.08, 19.4 min, 73 steps
GPT-6 AstraOpenAI72.8%mini-swe-agent, medium effort2026-09-22pass@4 82.3%, CI ±2.6, $3.08, 14.7 min, 26 steps
GPT-5.6 SolOpenAI72.7%mini-swe-agent, max effort2026-09-22pass@4 85.8%, CI ±2.8, $8.39, 18.8 min, 61 steps
Gemini 3.8 FlashGoogle71.0%mini-swe-agent, medium effort2026-09-22pass@4 83.2%, CI ±2.3, $1.97, 13.8 min, 147 steps
GPT-5.6 SolOpenAI70.7%mini-swe-agent, xhigh effort2026-09-22pass@4 85.8%, CI ±0.8, $4.70, 13.3 min, 44 steps
Claude Fable 5Anthropic69.9%mini-swe-agent, xhigh effort2026-09-22pass@4 88.5%, CI ±3.2, $13.41, 23.5 min, 68 steps
Claude Fable 5Anthropic69.7%mini-swe-agent, max effort2026-09-22pass@4 84.1%, CI ±4.0, $21.63, 34.9 min, 88 steps
GPT-5.6 TerraOpenAI69.6%mini-swe-agent, max effort2026-09-22pass@4 88.5%, CI ±2.6, $4.95, 16.9 min, 76 steps
GLM-5.3Z.ai69.0%mini-swe-agent, max effort2026-09-22pass@4 87.6%, CI ±3.0, $3.99, 35.3 min, 124 steps
Claude Opus 5Anthropic68.9%mini-swe-agent, medium effort2026-09-22pass@4 89.4%, CI ±1.2, $3.29, 12.7 min, 52 steps
Kimi K3Moonshot68.5%mini-swe-agent, max effort2026-09-22pass@4 89.4%, CI ±4.5, $4.65, 75.7 min, 98 steps
Grok 4.6xAI67.5%mini-swe-agent, medium effort2026-09-22pass@4 84.1%, CI ±2.3, $3.45, 14.9 min, 70 steps
GPT-5.6 LunaOpenAI67.2%mini-swe-agent, max effort2026-09-22pass@4 90.3%, CI ±4.0, $3.03, 18.7 min, 102 steps
GPT-5.5OpenAI67.0%mini-swe-agent, xhigh effort2026-09-22pass@4 88.5%, CI ±6.5, $7.23, 30.1 min, 82 steps
GPT-6 AstraOpenAI67.0%mini-swe-agent, low effort2026-09-22pass@4 79.6%, CI ±1.3, $1.60, 10.2 min, 20 steps
Gemini 3.7 FlashGoogle65.5%mini-swe-agent, medium effort2026-09-22pass@4 83.2%, CI ±3.1, $2.03, 20.7 min, 117 steps
GLM-5.3 FlashZ.ai63.4%mini-swe-agent, max effort2026-09-22pass@4 85.0%, CI ±4.4, $0.48, 25.6 min, 123 steps
DeepSeek V4 ProDeepSeek62.8%mini-swe-agent, max effort2026-09-22pass@4 88.5%, CI ±6.3, $0.24, 36.9 min, 155 steps
Claude Opus 4.8Anthropic59.0%mini-swe-agent, max effort2026-09-22pass@4 79.3%, CI ±1.8, $13.22, 58.2 min, 120 steps
Qwen3.8 MaxAlibaba57.5%mini-swe-agent, xhigh effort2026-09-22pass@4 83.2%, CI ±2.7, $3.73, 42.8 min, 111 steps
Muse Spark 1.2Meta54.9%mini-swe-agent, xhigh effort2026-09-22pass@4 81.4%, CI ±2.1, $3.70, 15.5 min, 101 steps
Claude Sonnet 5Anthropic53.8%mini-swe-agent, max effort2026-09-22pass@4 78.8%, CI ±4.2, $26.40, 80.1 min, 268 steps
DeepSeek V4 FlashDeepSeek53.3%mini-swe-agent, max effort2026-09-22pass@4 80.5%, CI ±3.6, $0.10, 24.0 min, 153 steps
GPT-5.4OpenAI51.8%mini-swe-agent, xhigh effort2026-09-22pass@4 77.9%, CI ±1.5, $5.65, 22.8 min, 70 steps
Gemini 3.1 Pro PreviewGoogle11.7%mini-swe-agent, high effort2026-09-22pass@4 28.3%, CI ±1.5, $2.14, 14.5 min, 76 steps
Datacurvemini-swe-agentv1.1, 113 taskspass@1deepswe.datacurve.ai

GPT-6 Astra at xhigh leads Claude Opus 5 at max by 0.5 points, inside both confidence intervals, so the order of the top five rows is not established. Gemini 3.1 Pro Preview at 11.7% sits more than 40 points below every other row. The data shows 452 scored attempts, 28,369 mean output tokens and 76 steps per attempt, and neither the data file, the blog nor the paper gives a cause. Lower-effort settings left out of the table include GPT-5.6 Sol high at 69.4%, Claude Fable 5 high at 68.6% and Grok 4.6 xhigh at 66.7%. Claude Opus 5.5, Claude Fable 5.1, GPT-6.1 Sol and Gemini 4 Argon are not on the board as of this snapshot.

DeepSWE v1.1: pass@1 vs cost per taskDatacurve's official board, mini-swe-agent, snapshot 2026-09-22
70%75%$2$5$10$20
FrontierOpenAIGoogleBehind the frontier
Cost per task, log scale, lower to the right
Source: Datacurve DeepSWE v1.1 leaderboard data, board snapshot 2026-09-22, accessed 2026-10-06. Cost is Datacurve's mean USD per attempt, computed from measured token counts at each provider's published list price.

Four points sit on the cost frontier: GPT-6 Astra low, Gemini 3.8 Flash medium, Gemini 3.8 Flash high and GPT-6 Astra xhigh. Every other configuration in the chart costs more for the same or lower score. Claude Opus 5 max scores 0.5 points below Astra xhigh at 2.7 times the cost. Three budget rows fall below the chart's range: DeepSeek V4 Pro max at 62.8% for $0.24, GLM-5.3 Flash max at 63.4% for $0.48 and DeepSeek V4 Flash max at 53.3% for $0.10.

Mercor independent run, v1.1, mini-swe-agent

Mercor ran the same 113 tasks in mini-swe-agent with three samples per task and its own settings; rows are comparable with each other and not with Datacurve's board or vendor numbers. Source is Mercor's DeepSWE v1.1 leaderboard, accessed 2026-10-06. The page lists 47 configurations, renders the top rows shown here, and gives no test date, step limit, time limit, cost or time per row.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic72.3%Mercor run, mini-swe-agent, maxAccessed 2026-10-06Error margin ±7.3, 336 samples
GPT-6.1 SolOpenAI72.3%Mercor run, mini-swe-agent, maxAccessed 2026-10-06Error margin ±7.4, 339 samples
GPT-6 AstraOpenAI72.0%Mercor run, mini-swe-agent, maxAccessed 2026-10-06Error margin ±7.5, 339 samples
DeepSeek V4.1 FlashDeepSeek71.7%Mercor run, mini-swe-agent, maxAccessed 2026-10-06Error margin ±6.3, 339 samples
Claude Opus 5Anthropic71.4%Mercor run, mini-swe-agent, maxAccessed 2026-10-06Error margin ±6.8, 339 samples
Gemini 3.8 FlashGoogle71.4%Mercor run, mini-swe-agent, highAccessed 2026-10-06Error margin ±6.8, 339 samples
Claude Opus 5Anthropic70.2%Mercor run, mini-swe-agent, xhighAccessed 2026-10-06Error margin ±6.9, 339 samples
Claude Fable 5.1Anthropic67.3%Mercor run, mini-swe-agent, highAccessed 2026-10-06Error margin ±6.9, 339 samples
Mercormini-swe-agentv1.1, 113 tasksmercor.com

Mercor's run is useful as a cross-check. Three configurations appear on both boards, and Mercor scores each a little lower: GPT-6 Astra max 72.0% against Datacurve's 73.2%, Claude Opus 5 max 71.4% against 73.6%, and Gemini 3.8 Flash high 71.4% against 73.8%. The top six Mercor rows sit within one point of each other, inside every error margin. Mercor is also the only independent source that has run Claude Opus 5.5, GPT-6.1 Sol and Claude Fable 5.1, and none of the three moves past GPT-6 Astra by more than 0.3 points.

Vendor-reported results

Each vendor ran its own model with its own harness and trial count; these are not comparable with Datacurve's board, Mercor's run or each other.

ModelOrganizationScoreHarness and setupDateSource and notes
Gemini 4 ArgonGoogle DeepMind77.9%Harness, effort and trials not published2026-09-30Gemini 4 Argon launch post, called "new state of the art"; methodology page returned 404
Claude Opus 5.5Anthropic74.2%Average over five trials; harness and effort not stated2026-09-22Claude Opus 5.5 System Card, section 8.3
v1.1

Google's 77.9% is the highest DeepSWE v1.1 number published anywhere, and it is also the least documented. Google has not published the harness, effort or trial count on the pages retrieved. Anthropic's 74.2% for Opus 5.5 is 1.9 points above the 72.3% Mercor measured for the same model. Datacurve's issue #98, opened 2026-09-22, notes that no Anthropic result for Opus 5.5 was on the official board yet.

Historical progression

The frozen v1 board used exit-code scoring. It is kept here as a record, and its scores do not line up with v1.1 scores for the same model. Source is the v1 leaderboard data, snapshot 2026-06-20.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-5.5OpenAI70.0%v1, mini-swe-agent, xhigh effort2026-06-20$6.61 per task
Claude Opus 4.8Anthropic58.2%v1, mini-swe-agent, max effort2026-06-20$12.58 per task
GPT-5.4OpenAI55.5%v1, mini-swe-agent, xhigh effort2026-06-20$4.38 per task
Claude Opus 4.7Anthropic54.2%v1, mini-swe-agent, max effort2026-06-20$18.19 per task
GLM-5.2Z.ai41.5%v1, mini-swe-agent, max effort2026-06-20$3.95 per task
Claude Sonnet 4.6Anthropic31.8%v1, mini-swe-agent, high effort2026-06-20$5.52 per task
Datacurvemini-swe-agentv1deepswe.datacurve.ai

The cleaner progression comes from the v1.1 board, where Datacurve has re-run older models next to new ones in the same harness. Each row is the best effort setting for that model; release dates are from Mercor's embedded data.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-5.4OpenAI51.8%v1.1, mini-swe-agent, xhigh effortReleased 2026-03-05Datacurve v1.1 board
GPT-5.5OpenAI67.0%v1.1, mini-swe-agent, xhigh effortReleased 2026-04-23+15.2 over GPT-5.4
GPT-5.6 SolOpenAI72.7%v1.1, mini-swe-agent, max effortReleased 2026-07-09+5.7 over GPT-5.5
GPT-6 AstraOpenAI74.1%v1.1, mini-swe-agent, xhigh effortReleased 2026-09-03+1.4 over GPT-5.6 Sol
Claude Opus 4.8Anthropic59.0%v1.1, mini-swe-agent, max effortReleased 2026-05-28Datacurve v1.1 board
Claude Opus 5Anthropic73.6%v1.1, mini-swe-agent, max effortReleased 2026-07-24+14.6 over Claude Opus 4.8
Datacurvemini-swe-agentv1.1deepswe.datacurve.ai

OpenAI's best configuration rose 22.3 points in six months, from GPT-5.4 to GPT-6 Astra. The gains are shrinking fast: 15.2, then 5.7, then 1.4 points. Anthropic's jump from Opus 4.8 to Opus 5 was 14.6 points. Mercor's run puts the same older models at GPT-5.4 xhigh 49.0% and Opus 4.8 max 59.6%, close to Datacurve's numbers but from a different run.

Failure modes and limitations

Most of what follows comes from the paper's HTML version. The one serious outside critique comes from Epoch AI, and it conflicts with the paper on a basic question.

Epoch AI rates v1.1 "Flawed." Epoch's review found issues in at least 23 of the 113 tasks, over 20.3%, which is above its own published threshold. The main cause is the clean-container grading that v1.1 introduced. The verifier applies the agent's patch in a fresh container and adds hidden tests. When the model and the hidden tests define the same names, compilation fails on symbol redeclaration. When agents put helpers in files the verifier deletes, they get undefined-reference errors. Epoch also lists underspecified requirements, such as CSS grid placement and plugin naming, and hidden patches that invalidate fixtures or change auto-generated test IDs. In Epoch's words, "Of the 23 observed instances of false negatives, 18 (78%) were due to an error that we believe could affect any task." That finding sits badly next to the paper's 1.4% disagreement rate, which came from Datacurve's own judge audit.

The verifier audit relies on an LLM judge. GPT-5.5 at xhigh effort judged 735 DeepSWE rollouts across 30 tasks and nine configurations. It found 2 false positives and 8 false negatives, 1.4% disagreement with a 95% CI of 0.7% to 2.5%. The same audit on 789 SWE-bench Pro rollouts found 67 false positives and 189 false negatives, 32.4%. The judge is one of the graded model families, and Datacurve has not released the judge prompt. Read the 1.4% as Datacurve's own measurement.

Contamination resistance has an expiry date. Tasks were never merged and agents get a shallow clone, so the fix is not in commit history or .git. Datacurve says itself that this is an evaluation-time property. The tasks and verifiers are now public, future training runs can ingest them, and the corpus has to be refreshed with new tasks to stay clean. The design is aimed at a documented problem. On SWE-bench Pro, 33 of 38 externally documented cheating cases involved recovering the merged fix from .git, and the paper's own audit flagged 18% of Opus 4.7 passes and 25% of Opus 4.6 passes there as recovering the gold solution.

Model families fail differently. Claude models most often miss one of several parallel requirements; about 67% of their "missed requirement" verdicts are cases where the agent shipped one branch of an enumerated list. GPT models have the lowest missed-requirement rate and the most stable reading of the prompt. Gemini 3 Flash skipped all verification on 18% of runs. Strong models write their own tests unprompted in over 80% of DeepSWE runs, against 20% to 40% on SWE-bench Pro, which the paper attributes to SWE-bench Pro's prompt forbidding test edits.

Pass/fail scoring hides near misses. A patch that meets 9 of 10 requirements fails. The verifiers check function only, so code quality, style, error handling and performance go unmeasured. Planning, review and explanation are out of scope, and ambiguous requests were excluded on purpose.

The noise is real at the top. The run-to-run 95% CI is ±0.8 to ±6.5 points per configuration, and it covers re-run noise only, not the noise from which 113 tasks were sampled. The top five rows span 74.1% to 73.2%. Anyone who calls one of them the winner is reading noise.

Cost does not track accuracy. The paper finds no consistent relationship between median cost, tokens or wall-clock time and pass rate. GPT-6 Astra at max costs $7.50 and scores 73.2%. At xhigh it costs $4.43 and scores 74.1%.

BenchmarkTasksGradingWhat it measures
DeepSWE v1.1113 from-scratch tasks, 91 repos, 5 languagesHand-written functional verifiers, binary pass@1Feature work in existing repositories, one fixed bash harness
SWE-bench ProMined from merged PRs, 11 public-split reposInherited testsRepository fixes drawn from public history
Terminal-Bench 4.066 containerized CLI tasksTask checksTerminal command sequencing in biology, physics, CAD, proofs and GPU work
FrontierCode150 tasks from real PRsBlocking functional criteria plus a weighted code-quality rubric, mean@5Function and code quality in Claude Code and Codex CLI
FrontierSWE v2Multi-hour challengesCovered on the FrontierSWE v2 pageLong autonomous engineering sessions

SWE-bench Pro (arXiv 2509.16941) is the benchmark DeepSWE positions itself against. Its prompts run 4,000-plus characters, twice DeepSWE's, and its tasks come from history that is already public. Anthropic's Opus 5.5 system card reports 89.9% on SWE-bench Pro and 74.2% on DeepSWE v1.1 for the same model. Those are different benchmarks, and the 15.7-point gap says nothing about a regression. It does show how much headroom DeepSWE still has where SWE-bench Pro is close to saturated for this model.

Terminal-Bench 4.0 tests terminal command sequencing on tasks such as computational biology, physics simulation, CAD, formal proofs and GPU work. It is the better signal if your agent mostly drives a shell. DeepSWE is the better signal for feature work inside an existing repository.

FrontierCode from Cognition grades code quality as well as function and runs in product harnesses. If you care whether a patch is clean enough to merge, FrontierCode measures that and DeepSWE does not. If you want to compare models with the scaffold held fixed, DeepSWE does that and FrontierCode does not.

FrontierSWE v2 runs for hours per challenge. DeepSWE attempts finish in 10 to 80 minutes on average, so it measures sustained work across one feature, not a day of autonomous engineering. Other neighbors worth reading are CursorBench, ProgramBench, BigCodeBench and the older function-level HumanEval. SWE-bench Multilingual, 300 problems in 9 languages, and SWE-bench Multimodal also appear in the Opus 5.5 system card, and both are mined from existing PRs like the rest of the SWE-bench family. For how all of these fit together, see the LLM benchmarks and LLM evaluation overviews.

What DeepSWE means for teams choosing a model

Price the effort level, not the model. Effort is where test-time compute shows up on the bill. GPT-6 Astra gains 5.8 points going from low ($1.60, 67.0%) to medium ($3.08, 72.8%), then only 1.3 more from medium to xhigh ($4.43, 74.1%). Max costs $7.50 and scores lower, at 73.2%. Claude Opus 5 climbs from medium ($3.29, 68.9%) to max ($11.84, 73.6%), 4.7 points for 3.6 times the cost. Astra medium is the sensible default and xhigh is the ceiling. Skip Astra max.

Gemini 3.8 Flash at high effort is the cost leader in the top tier. It scores 73.8% at $2.36 and 11.4 minutes per task. Astra xhigh scores 74.1% at $4.43 and 18.9 minutes, and Opus 5 max scores 73.6% at $11.84 and 31.9 minutes. Watch the step count. Gemini takes 166 steps per task against Astra's 29, so if your agent loop is bound by per-step latency or rate limits, Astra's fewer, heavier steps fit better.

If you run retries or best-of-N, look at pass@4. Opus 5 max solves 88.5% of tasks in at least one of four attempts and Gemini 3.8 Flash high solves 85.8%, against 80.5% for Astra xhigh. Astra is more consistent per attempt. Opus and Gemini cover more tasks when you can afford to sample and verify.

Fable-tier prices don't pay on this kind of work. Claude Fable 5 at xhigh scores 69.9% at $13.41, 4.2 points below Astra xhigh at three times the cost. Fable 5 max costs $21.63 for 69.7%.

Open-weight and budget models trail by 10 to 12 points at a ninth to an eighteenth of the cost. DeepSeek V4 Pro max scores 62.8% at $0.24 and GLM-5.3 Flash max 63.4% at $0.48, against 74.1% at $4.43 for the leader. That trade works for first-pass or low-stakes tasks where a failed attempt gets retried. DeepSeek V4 Pro's ±6.3 CI is wide, so test it on your own tasks before committing.

Choose from Datacurve's board and treat vendor headlines as claims. Neither Argon's 77.9% nor Opus 5.5's 74.2% came from the Datacurve harness, and neither model is on that board. Mercor's independent run puts Opus 5.5 and GPT-6.1 Sol both at 72.3%, level with GPT-6 Astra's 72.0% on the same run. Nothing published so far shows any model clearly ahead of GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5 under the same conditions. Given Epoch's finding of problems in at least one task in five, a gap of a point or two at the top is not a reason to switch.

The Klu LLM leaderboard tracks these models across other benchmarks. To run the same kind of pass-rate and cost-per-task eval on your own repositories and tickets, see how engineering teams set it up in Klu on the engineering page.