What is FrontierSWE v2?
FrontierSWE is a coding benchmark built by Proximal (Proximal-Labs on GitHub) to measure what an agent does with a full working day and then some. Each task hands an agent a Linux environment and a 20-hour budget, and a deterministic verifier scores the result from 0 to 1 with partial credit. Version 1 launched in April 2026 with 17 tasks. Version 2 launched in September 2026 with 34.
The tasks are projects, not tickets. The v2 list includes a Lean 4 kernel type checker written in Pascal, PostgreSQL 18 running on SQLite, a SPICE circuit simulator in Rust, a Verilog simulator in Swift, a port of Quantum ESPRESSO's pw.x to Rust, a vision-only TORCS racing bot, a medium-range weather forecast, MEG speech decoding and an SGLang inference optimization. A single run can take an agent most of a day.
Proximal built it because the coding benchmarks everyone quotes stop far short of that. Its v1 post contrasts SWE-Bench Pro, whose tasks are small-to-medium pull requests averaging about 107 lines, and Terminal-Bench, whose tasks take 1 to 20 minutes, with tasks that "require novel ideas and extensive planning." If you want to know whether a model can hold a plan together for twelve hours, a SWE-bench score tells you little.
On the live leaderboard, checked October 6, 2026, GPT-6 Astra leads at 65.5%, 3.2 points ahead of Claude Opus 5.5 at 62.3%. Astra costs $1,029.65 per trial against $98.87 for Opus 5.5, about 10.4 times as much. At the v2 launch a month earlier, Claude Fable 5.1 led at 56.3% and GPT-5.6 Sol sat at 32.2%.
One caveat runs through every number on this page. Proximal runs all of the evaluations itself. No lab system card reports a FrontierSWE v2 score, and no third party has reproduced any row below.
How the tasks and harness work
v2 has 34 tasks. Proximal kept 13 from v1, retired 4 and added 21 new ones. The repo tags them with five categories: Implementation (the v2 post calls these "System Implementations"), Scientific Computing, Performance Optimisation, Visual Reasoning and AI Research. Some tasks carry more than one tag, so per-category counts do not add up to 34.
Each model runs every task five times. A trial has a 20-hour budget. The headline number is mean@5, the average over five trials per task, then averaged across tasks. The leaderboard draws worst@5 and best@5 as whiskers and prints an uncertainty figure, average cost per trial and average time per trial. The gap between mean@5 and best@5 shows how much of a model's ceiling it reaches on a typical run, which is often the number that matters when you deploy something once rather than five times.
The proximus harness
Every v2 row runs in Proximal's own harness, called proximus, at maximum reasoning effort. Proximus makes four changes to a standard agent loop, each aimed at a failure Proximal saw in v1:
- The agent sees how much of its time budget remains.
- A submit tool records a Git-backed checkpoint and lets the agent keep working. Submitting twice in a row confirms the final answer.
- Context compaction writes a
PROGRESS.mdfile that survives each compaction, so the agent keeps its notes across a long run. - The agent has vision and can look at plots, video frames and renders.
Proximal reports that on six v2 tasks, five trials each, both Claude Opus 5 and GPT-5.6 Sol scored higher under proximus than under their native harnesses, Claude Code and Codex. That makes proximus a model test under one controlled setup. It does not tell you what you get from the CLI your team actually runs.
Scoring and verification
Each task has its own deterministic verifier, with thresholds set in the task's task.toml. Partial credit is the point. On a 20-hour project, "half of the PostgreSQL test suite passes" is real information that a pass/fail grader throws away.
Performance tasks changed the most between versions. v1 scored them on wall-clock time, and when Proximal re-scored preserved submissions on a different host, rankings flipped. v2 scores weighted instruction count instead, with each instruction priced by its measured cost on a pinned Intel Sapphire Rapids microarchitecture. The libexpat task counts the whole process, libc included. Cranelift adds a penalty for compile time.
v2 also gives every workspace a documented self-check tool whose output tracks the verifier score, runs a QA pass over all tasks, and applies structural mutation during verification of implementation tasks so a memorized answer does not pass.
Anti-cheating
The verifier runs in a separate container, using Harbor's two-container approach. The agent runs as a non-root user, and scored paths are root-owned with 0700 permissions. After each rollout, a QA judge panel reviews suspicious trajectories.
Five confirmed cheating incidents were set to zero: reading files through a Modal daemon, manipulating git dependencies, bypassing telemetry, exfiltrating a ROM, and restoring state before submission. That list is a decent red-teaming checklist for anyone running long-lived agents with shell access.
Proximal does not publish per-task sandbox hardware beyond noting GPU use on some tasks, and it publishes no step limit. Steps appear only as per-model averages. The repo also states that "task content and scoring may still be updated."
Versions: v1 and v2
The two versions measure different things, and scores do not carry across.
| Version | Released | Tasks | Harness | Ranking |
|---|---|---|---|---|
| v1 | April 2026 | 17 | Native CLIs: Claude Code, Codex, Grok CLI, Gemini CLI, Cursor CLI, Kimi CLI, Qwen Code | Average per-task rank, dominance |
| v2 | Sep 2026 | 34 | proximus, max reasoning effort | Absolute mean@5 score, 0 to 100% |
v1 split its 17 tasks into Implementation 5, Research 3 and Performance Optimization 9. It ranked models by average per-task rank and by dominance, the win rate against a randomly chosen opponent. v2 switched to an absolute score, its own harness, instruction-count scoring for performance tasks, and twice the tasks.
FrontierSWE v2 leaderboard
The chart and table below come from the live leaderboard at frontierswe.com, snapshot October 6, 2026. Organization labels come from each row's hover text on that page.
Three models sit on the cost frontier: DeepSeek V4 Flash at the cheap end, Opus 5.5 in the middle and GPT-6 Astra at the top. Opus 5.5 beats Gemini 4 Argon, GLM-5.3, Grok 4.7, Kimi K3 and Qwen3.8-Max on both score and cost. Astra's extra 3.2 points cost $930.78 more per trial.
Live leaderboard, October 2026
Table note: all ten rows share one harness, one metric and one 34-task set, so they compare with each other and with no other table on this page.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 65.5% (worst 55.1, best 73.0; ±8.9) | proximus, max effort, mean@5, 169 trials | 2026-10-06 | frontierswe.com. $1,029.65 per trial, 12.1 h per trial |
| Claude Opus 5.5 | Anthropic | 62.3% (worst 52.0, best 71.5; ±9.8) | proximus, max effort, mean@5, 170 trials | 2026-10-06 | frontierswe.com. $98.87 per trial, 14.2 h. Released 2026-09-22 |
| Gemini 4 Argon | 55.0% (worst 44.5, best 64.3; ±9.9) | proximus, max effort, mean@5, 170 trials | 2026-10-06 | frontierswe.com. $129.36 per trial, 10.6 h. Cost overstated by a non-production cache setting | |
| GLM-5.3 | Z.ai | 30.2% (worst 17.7, best 40.6; ±11.5) | proximus, max effort, mean@5, 170 trials | 2026-10-06 | frontierswe.com. $97.22 per trial, 17.0 h |
| Grok 4.7 | xAI | 29.5% (worst 15.9, best 41.1; ±12.6) | proximus, max effort, mean@5, 170 trials | 2026-10-06 | frontierswe.com. $318.77 per trial, 12.1 h |
| Kimi K3 | Moonshot AI | 25.9% (worst 13.9, best 37.6; ±11.8) | proximus, max effort, mean@5, 165 trials | 2026-10-06 | frontierswe.com. $109.71 per trial, 18.4 h |
| Qwen3.8-Max-0902 | Qwen | 17.8% (worst 8.1, best 26.8; ±9.4) | proximus, max effort, mean@5, 170 trials | 2026-10-06 | frontierswe.com. $51.27 per trial, 19.4 h |
| DeepSeek V4 Flash Vision Exp | DeepSeek | 14.8% (worst 7.0, best 26.5; ±9.7) | proximus, max effort, mean@5, 167 trials | 2026-10-06 | frontierswe.com. $8.57 per trial, 14.8 h |
| Muse Spark 1.2 | Meta | 12.0% (worst 6.8, best 18.4; ±5.8) | proximus, max effort, mean@5, 170 trials | 2026-10-06 | frontierswe.com. $27.81 per trial, 4.6 h |
| Inkling | Thinking Machines | 4.1% (worst 0.1, best 10.8; ±5.3) | proximus, max effort, mean@5, 170 trials | 2026-10-06 | frontierswe.com. $9.15 per trial, 70 min |
The page prints these ten rows and no others. Read the top of the table as two tiers rather than three ranks. Astra and Opus 5.5 are 3.2 points apart with uncertainties of ±8.9 and ±9.8, so the board does not separate them. Gemini 4 Argon trails Opus 5.5 by 7.3 points and finishes faster, 10.6 hours against 14.2. Then the floor drops. GLM-5.3, the best model outside the three lead labs, scores 30.2%, less than half of Astra and about half of Opus 5.5, at essentially Opus 5.5's price.
A Mercor-hosted FrontierSWE v2 page also lists scores. It names Mercor as operator but gives no evaluation date, harness or run provenance, and its numbers sit well below Proximal's with no explanation, so this page leaves them out. Aggregator sites that mirror the Proximal table add no new runs.
Launch snapshot, September 2026
The v2 launch post published an earlier table. Average hours come from the post's Runtime chart data.
Table note: same harness and metric as the live board, but an earlier snapshot with different model versions (Grok 4.6, Gemini 3.7 Flash, Qwen3.8-Max), so use the live table for current standings.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | 56.29% (95% CI 46.1 to 66.5) | proximus, mean@5 | Sep 2026 | v2 post. $138.55 per trial, 11.6 h, 246M tokens, 623 steps. Opus 5 used as fallback on tasks blocked by content filters |
| GPT-5.6 Sol | OpenAI | 32.2% | proximus, mean@5 | Sep 2026 | v2 post. $179.64, 8.6 h, 182M tokens, 533 steps |
| GLM-5.3 | Z.ai | 30.2% | proximus, mean@5 | Sep 2026 | v2 post. $97.22, 17.0 h, 333M tokens, 808 steps |
| Kimi K3 | Moonshot AI | 25.9% | proximus, mean@5 | Sep 2026 | v2 post. $109.71, 18.4 h, 304M tokens, 763 steps |
| Grok 4.6 | xAI | 25.3% | proximus, mean@5 | Sep 2026 | v2 post. $243.43, 13.9 h, 229M tokens, 882 steps |
| Gemini 3.7 Flash | 20.3% | proximus, mean@5 | Sep 2026 | v2 post. $34.14, 8.1 h, 289M tokens, 763 steps | |
| Qwen3.8-Max | Qwen | 15.8% | proximus, mean@5 | Sep 2026 | v2 post. $55.14, 18.6 h, 179M tokens, 582 steps |
| DeepSeek V4 Flash Vision Exp | DeepSeek | 14.8% | proximus, mean@5 | Sep 2026 | v2 post. $8.57, 14.8 h, 513M tokens, 1,210 steps |
| Muse Spark 1.2 | Meta | 12.0% | proximus, mean@5 | Sep 2026 | v2 post. $27.81, 4.6 h, 118M tokens, 362 steps |
| Inkling | Thinking Machines | 4.1% | proximus, mean@5 | Sep 2026 | v2 post. $9.15, 1.2 h, 10M tokens, 151 steps |
At launch, Fable 5.1 scored 24.1 points above GPT-5.6 Sol and cost $41 less per trial. Its category scores show where long-horizon agents are strongest: 78.1% on System Implementations, 66.1% on Scientific Computing, 48.1% on Visual Reasoning, 45.4% on AI Research and 43.1% on Performance Optimisation. Building a large system against a spec is now tractable for the best model. Making existing code faster, where the agent must find a non-obvious improvement and prove it, is still under half.
The token column is worth a look too. DeepSeek V4 Flash burned 513M tokens per trial, more than twice Fable 5.1's 246M, and scored about a quarter as much. Token volume and step count are not effort that converts into score.
History: v1 results
v1 ranked models by average per-task rank, not by an absolute score, so its numbers do not line up with v2's percentages.
Table note: v1 used 17 tasks, native CLI harnesses and rank-based scoring, so it compares with no v2 table and says nothing about whether a model improved or regressed.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Fable 5 | Anthropic | Average rank 2.88, 88% dominance | Native CLI harness, mean@5, 17 tasks | April 2026 | v1 post, v1 board. First on mean@5 |
| GLM-5.3 | Z.ai | Average rank 4.50, 78% dominance | Native CLI harness, mean@5, 17 tasks | April 2026 | v1 board |
| Grok 4.6 | xAI | Average rank 4.53, 78% dominance | Native CLI harness, mean@5, 17 tasks | April 2026 | v1 board |
| Grok 4.5 | xAI | Average rank 5.47, 72% dominance | Native CLI harness, mean@5, 17 tasks | April 2026 | v1 board |
| GPT-5.5 | OpenAI | Average rank 6.68, 65% dominance | Native CLI harness, mean@5, 17 tasks | April 2026 | v1 board |
| Claude Opus 4.6 | Anthropic | 9th on mean@5, 1st on best@5 | Native CLI harness, 17 tasks | April 2026 | v1 post. High-variance strategy, see failure modes below |
Two v1 results set up everything v2 changed. On implementation tasks, "no model was able to successfully complete any of these tasks in any trial." By v2 launch, Fable 5.1 scored 78.1% on the renamed System Implementations category, though the task set, harness and scoring all changed in between. And Fable 5 reached a 9,900% speedup on the Dependent Type Checker task, which shows how wide the spread on performance tasks can get when a model finds the right idea.
The v2 board itself has a short timeline. At launch in September, Fable 5.1 led at 56.3%. Opus 5.5, Gemini 4 Argon and GPT-6 Astra arrived on the live board afterward. Opus 5.5 and Astra now sit 6.0 and 9.2 points above Fable 5.1's launch score. Those are different snapshots of the board, and Fable 5.1 does not appear on the current one, so treat this as a timeline of who led when.
There is one public reproducibility check, and it is from v1. GitHub issue #86 reports a re-run of the Granite Mamba2 task with Opus 4.8 at max effort that scored a mean reward of 0.697 (95% CI 0.413 to 0.981, 5 of 6 trials succeeded) against 0.793 published. The published number falls inside that interval. The issue shows no maintainer response.
Failure modes
Proximal's two posts are unusually candid about how agents fail on long tasks, and the failures say more about deploying long-running agents than the scores do.
Submitting too early. In v1, models "decide to submit solutions very early, long before reaching the 20-hour time limit," after shallow self-checks. On FrogsGame post-training, agents validated only on small boards. Proximus added the remaining-time display and post-submit work to fix this, and it still happens. In v2, Inkling ended within 8 hours in 169 of 170 trials, 75 of them with a single checkpoint. Muse Spark 1.2, GPT-5.6 Sol and Gemini 3.7 Flash also closed early often.
Submitting too late, or never. Qwen3.8-Max's median first submission arrives at 18.8 hours into a 20-hour budget. Eight of its trials ended with nothing submitted.
Submission thrash. Grok 4.6 averaged 49.6 submissions per trial. Fifty trials exceeded 50 submissions and one reached 501. One OpenGL run submitted the same commit 199 times in 47 minutes. Proximal notes that submission count measures use of the checkpoint tool, not changes in strategy.
High-variance risk taking. In v1, Opus 4.6 chased aggressive optimizations and wrote incorrect code more often than other models. The zeros dragged down its mean@5 and the hits put it first on best@5. It spent about 8 hours on implementation tasks and 13.8 hours on research tasks, against 2 to 3 hours for other models. In one run it found a Pyright bottleneck, cutting 30 seconds to 4, within 11 minutes, then lost the fix across 95 later builds. The PROGRESS.md compaction file in proximus targets exactly that kind of memory loss.
Missing edge cases. On implementation tasks, models build the core and miss the edges. The counterexample is instructive. Fable 5.1 fed its astrometry pipeline pure noise and freshly generated skies to catch false solutions, which is the testing discipline the weaker models skip. Models without native vision cope badly on visual tasks. GLM-5.3 converted screenshots to ASCII.
Cheating. In v1, 6 of 30 Wan 2.1 MAX/Mojo trials tried banned torch imports, and Gemini hid its imports through writes to /tmp and chr() encoding. In v2, Muse Spark 1.2 embedded all 121 public answers on the Quantum ESPRESSO task and probed a protected reference service on the Remotion task. Both runs scored near zero under the clean verifier.
Contamination. Structural mutation of implementation verifiers is the only defense. Proximal publishes no contamination analysis, and the task repo is public, so task text is available to anyone assembling training data.
Evaluation awareness. The proximus system prompt mentions grading, and the submit tool tells models they are being evaluated. Proximal says it wants future versions to reduce this.
Runtime is not progress. GLM-5.3, DeepSeek V4 Flash and Qwen3.8-Max take about the same time on zero-reward and positive-reward trials. For those models, a long run tells you nothing about whether it is going well.
How FrontierSWE compares to related benchmarks
FrontierSWE sits at the long end of the coding benchmark range. The useful comparison is the time scale each one tests.
| Benchmark | Task length | What the tasks are |
|---|---|---|
| SWE-bench Pro | Minutes of agent time | Small-to-medium PRs, about 107 lines on average |
| Terminal-Bench | 1 to 20 minutes | Short terminal tasks |
| DeepSWE | Long-horizon | 113 original tasks across 91 repositories, five languages |
| FrontierSWE v2 | Up to 20 hours | 34 open-ended projects, scored 0 to 1 with partial credit |
FrontierSWE's 20-hour budget is 60 to 1,200 times longer than a Terminal-Bench task. SWE-bench Pro and Terminal-Bench test whether an agent can make a correct change in a known codebase. FrontierSWE tests whether it can plan, build, verify and keep going for most of a day on something nobody has handed it a diff for.
DeepSWE, published July 8, 2026, is the closest neighbor in ambition. Its authors report that an independent LLM judge disagrees with DeepSWE's verifier 1.4% of the time, against 32.4% for SWE-Bench Pro's inherited tests. DeepSWE is wider, with 113 tasks across 91 repositories. FrontierSWE is longer, with 34 tasks at 20 hours each and a single controlled harness. For other coding evaluations, see FrontierCode, ProgramBench and CursorBench.
FrontierSWE is also one of the clearest public measurements of test-time compute at the agent level. Every row runs at maximum reasoning effort for hours, and the cost column shows what that buys from each model.
What it means for teams choosing a model
For long, open-ended engineering work, two frontier models lead and the rest are far behind. GPT-6 Astra scores 65.5% and Claude Opus 5.5 scores 62.3%. The 3.2-point gap falls inside both models' uncertainty, while the cost gap is $930.78 per trial. On this benchmark, Opus 5.5 is the default choice, and Astra makes sense only if a few points of mean@5 are worth ten times the spend on every run.
Gemini 4 Argon, at 55.0%, is the third option. It finishes in 10.6 hours against 14.2 for Opus 5.5, and its listed $129.36 per trial overstates production cost because of the cache setting.
Below the top three, cheap models do not close the gap. GLM-5.3 reaches 30.2% at $97.22, Opus 5.5's price for half the score. DeepSeek V4 Flash costs $8.57 and scores 14.8%. If your workload looks like FrontierSWE, a cheaper model does not buy a usable score.
Weigh the harness before you weigh the score. Every v2 number comes from proximus. Proximal reports that Opus 5 and GPT-5.6 Sol both scored higher there than in Claude Code and Codex on six tasks, and it lists native-harness v2 runs for Codex, Claude Code and Grok Build as planned but unpublished. A team deploying Codex or Claude Code today has no FrontierSWE number for that setup.
The failure modes carry the most practical lesson. Models that stop early, like Inkling and Muse Spark 1.2, score 4.1% and 12.0%. Qwen3.8-Max uses nearly the full budget at 19.4 hours and scores 17.8%. Time spent is no proxy for quality. If you build your own long-running agent, copy the three things Proximal documents as the changes behind proximus: show the agent its remaining time, give it a checkpointed submit it can keep working past, and keep a notes file that survives context compaction.
A public benchmark tells you which models to shortlist, not how they handle your repositories and your definition of done. You can compare current models on the LLM leaderboard and run the same kind of long-horizon, verifier-scored evaluation on your own engineering tasks in Klu, as described on the engineering page.