FrontierCode

Cognition's coding benchmark that asks whether an agent's patch to a real open-source repository is one the project's maintainers would merge, graded on six rubric axes

Agentic coding

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 15 min read

What is FrontierCode 1.1?

FrontierCode is Cognition's coding benchmark, and it asks a harder question than most. A patch passing the tests is not enough. The patch has to be one the repository's maintainers would merge. Cognition published version 1.0 on June 8, 2026, version 1.1 on July 7, and opened the public leaderboard on July 17.

The benchmark has 150 tasks drawn from 36 flagship open-source repositories. More than 20 maintainers of those projects wrote the tasks, and each task took more than 40 hours to build. Extended is all 150 tasks. Main is the 100 hardest. A 50-task Diamond subset existed in 1.0 and is gone in 1.1.

Cognition built it as a response to SWE-bench Verified and SWE-Bench Pro, which grade functional correctness and nothing else. Their tests fail in two directions. Incomplete tests accept wrong solutions, and over-specific tests reject valid ones. Cognition cites a METR note from March 10, 2026 finding that many SWE-bench-passing PRs would not be merged. Against SWE-Bench Pro, Cognition reports an 81% lower false positive rate, prompts one third the length, and three times as many languages.

The headline number is simple to state. On FrontierCode 1.1 Main, Claude Opus 5.5 at medium effort in Claude Code scores 54.6%, and the best non-Anthropic model, GPT-6 Astra at max effort in Codex, scores 53.3%. The gap in cost is far larger than the gap in score.

How the tasks and grading work

Tasks written by maintainers

Maintainers hand-picked each task from multi-PR chains and freeform requests rather than scraping single PRs. A prompt is a short, human-sounding task description plus the repository's guidelines for testing, lint and style, the kind of thing you would put in an AGENTS.md file. The agent has to work out what the maintainer wants. Cognition says FrontierCode patches are smaller than DeepSWE patches and the benchmark is still harder, because difficulty comes from how strict the rubric is, not from how many lines change.

Tasks are private. Cognition does not release them, to keep them out of training data, and it is opening evaluation to all model creators instead.

Six grading axes

Every task has a rubric that scores the patch on six axes:

  • Behavioral correctness: does it do what was asked.
  • Regression safety: does it break anything else.
  • Mechanical cleanliness: build, lint and style checks.
  • Test correctness: are the agent's own tests right.
  • Scope: does it touch only what it needs to.
  • Code quality: conventions, design patterns and readability.

Scope is the axis that catches the most capable models off guard. A model that helpfully refactors a neighboring function loses points even if the refactor is good. That shows up later in the effort results.

Grader types

Rubric items are checked by six kinds of grader. Classical graders inject test files, run them and clean up. Command graders pass when a shell command exits 0. Reverse-classical graders require that the agent's own tests fail on the base commit, which proves the tests exercise the change. Scope graders check file boundaries, diff size and, optionally, semantic locality. Prompt graders have an LLM review the diff against a natural-language rubric, with a threshold to meet, which makes part of the score LLM-as-a-judge.

The sixth type is the interesting one. Adaptive classical grading uses an LLM tool Cognition calls mutagent to patch the test files or the app code so they match the agent's implementation. That lets a valid solution with a different function signature or file layout pass, which is the false-negative problem SWE-bench never solved.

Blockers, score and pass rate

Each rubric criterion is either a blocker or a non-blocker. A blocker is something that would stop a code review cold, including performance or scope limits. A non-blocker covers style, type safety and readability. Fail any blocker and the run scores 0.

The leaderboard reports two numbers. Score is the weighted aggregate of rubric items, with blocker failures and runs flagged for unfair internet use set to zero. Pass rate is the fraction of runs that clear every blocker. Score always sits below pass rate in Cognition's data, because a run can clear the blockers and still drop non-blocker points. Opus 5.5 on Main scores 54.6% with a 59.6% pass rate.

Each model runs five times at every reasoning effort it offers. The metric is the five-trial average, and the leaderboard ranks each model at its best-scoring effort.

Quality control

Subjective rubrics need checking, and Cognition's process has several layers. The task author documents each rubric item and writes a "hack report," a lazy or adversarial wrong solution to test for false positives and a valid alternative to test for false negatives. Devin then tries to hack the rubric. The author writes four reference solutions spanning 0 to 100% to calibrate the scoring. A pod lead reviews the task, and a Cognition researcher reviews every task after that.

Internet access

Internet stays on in 1.1, because some tasks need API lookups and models lean on search. The prompt defines fair use (documentation, API references, error messages) and unfair use (anything that reveals the solution). A programmatic verifier flags references to the source PRs, upstream patches or files, and mirrors or vendored copies that contain the solution. Flagged runs score 0.

Cognition tried a domain blocklist first. It grew to about 1,200 domains, agents kept finding workarounds, and some spent more than 20 turns fighting it, so Cognition dropped it. An allowlist would need a custom list for every task. Since August 6 the leaderboard has shown each model's flag rate.

Harness per model

Each model runs in its vendor's own agent. Claude models run in Claude Code, GPT models in Codex, SWE-2 in Devin, Grok in grok-build, Gemini and SWE-1.x in chisel, and Composer in Cursor CLI. Kimi, GLM, MiniMax, Qwen and Inkling run in mini-swe-agent. A leaderboard row is a model, a harness and an effort setting together.

Cognition publishes no per-task timeout, step cap, tool allowlist, container spec or confidence interval. The data file does report tokens, cost, tool calls and steps per rollout, and minutes per rollout for some configurations.

Versions: 1.0 and 1.1

Version 1.1 made three changes. It added the fair-use prompt and the internet verifier. It audited all 1,000+ blocker criteria and demoted 75 overly strict ones to non-blockers. And it dropped Diamond, because the 50 tasks no longer reflected the 50 hardest and their low solve rates made the subset noisy.

The audit moved absolute scores. Cognition says model ordering stayed the same between versions, so a 1.0 score and a 1.1 score are different measurements even for the same model and harness.

VersionReleasedSubsetsWhat changed
1.02026-06-08Diamond (50), Main (100), Extended (150)Launch
1.12026-07-07Main (100), Extended (150)Fair-use prompt and verifier, 75 blockers demoted, Diamond dropped; absolute scores changed

FrontierCode 1.1 leaderboard

Cognition runs every score itself. Anthropic's launch pages republish the same Main numbers, and nobody outside Cognition has re-run the benchmark. The board lists 42 models as of its September 29 changelog entry, which added GPT-6.1 Sol. All figures below come from Cognition's leaderboard data file, the file the leaderboard page loads, accessed October 6, 2026.

FrontierCode 1.1 Main: score vs cost per task100 tasks, 5-trial average, each model in its vendor harness
45%50%55%$0.1$0.2$0.5$1$2$5$10
FrontierOpenAIAnthropicBehind the frontier
Cost per task, log scale, lower to the right
Source: Cognition's FrontierCode leaderboard data file (cognition.com/data/frontiercode-leaderboard/data.json), accessed 2026-10-06. Cost is mean USD per rollout as Cognition reports it, priced by Cognition, with Astra, Terra and Luna pricing corrected on 2026-09-10. Each point is one model at one reasoning effort.

Three of the ten charted points sit on the cost frontier, meaning no other point both scores higher and costs less. GPT-6 Luna at max scores 42.4% for $0.10. GPT-6.1 Sol at medium scores 50.2% for $0.36. Opus 5.5 at medium scores 54.6% for $0.80. Every other point costs more than $0.80 and scores lower than 54.6%. Across all 125 configurations in the data file, the frontier above 40% is the same three points plus GPT-6.1 Sol at low effort (45.5%, $0.25).

Opus 5.5 medium costs 17% of what GPT-6 Astra max costs and scores 1.3 points higher. Fable 5 at xhigh costs 16 times as much as Opus 5.5 medium and scores 1.1 points lower. Sonnet 5.5 at high effort (49.4%, $0.42) misses the frontier by a hair, because GPT-6.1 Sol medium scores 0.8 points more for $0.36. A time chart isn't possible, since Cognition publishes minutes per task for only Opus 5.5, Sonnet 5.5 and GPT-6.1 Sol.

Main, 100 tasks, best effort per model

Comparable to: other rows in this table (same version and subset, maintainer-run, harness and effort differ per row). Not comparable to Extended, the Devin Fusion runs or any 1.0 score.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic54.6% (pass rate 59.6%)Claude Code, medium, 5 trialsAdded 2026-09-22Cognition data. $0.80 per rollout, flag rate 0.0%
Claude Fable 5Anthropic53.5% (pass rate 58.9%)Claude Code, xhigh, 5 trialsBoard 2026-10-06Cognition data. $13.09 per rollout, flag rate 0.3%
Claude Opus 5Anthropic53.4% (pass rate 58.9%)Claude Code, medium, 5 trialsAdded 2026-07-24Cognition data. $4.31 per rollout, flag rate 0.6%
GPT-6 AstraOpenAI53.3% (pass rate 58.8%)Codex, max, 5 trialsAdded 2026-09-03Cognition data. $4.59 per rollout after the 2026-09-10 pricing correction, flag rate 0.0%
Claude Sonnet 5.5Anthropic52.1% (pass rate 57.2%)Claude Code, xhigh, 5 trialsAdded 2026-09-28Cognition data. $1.59 per rollout, 12.6 min, flag rate 0.0%
Claude Fable 5.1Anthropic50.9% (pass rate 55.5%)Claude Code, medium, 5 trialsAdded 2026-09-01Cognition data. $3.28 per rollout, flag rate 0.0%. Anthropic plots Fable 5.1 Low at 52.8%; see below
GPT-6.1 SolOpenAI50.2% (pass rate 55.8%)Codex, medium, 5 trialsAdded 2026-09-29Cognition data. $0.36 per rollout, 14.4 min, flag rate 0.4%
SWE-2Cognition50.0% (pass rate 55.5%)Devin, max, 5 trialsAdded 2026-09-10Cognition data. $1.18 per rollout, flag rate 0.0%. Cognition's own model on Cognition's benchmark
GPT-6 SolOpenAI49.3% (pass rate 54.3%)Codex, max, 5 trialsAdded 2026-09-22Cognition data. $2.07 per rollout, flag rate 0.0%
Grok 4.6xAI48.0% (pass rate 53.1%)grok-build, high, 5 trialsBoard 2026-10-06Cognition data. $2.88 per rollout, flag rate 0.7%
Grok 4.7xAI47.6% (pass rate 53.1%)grok-build, high, 5 trialsBoard 2026-10-06Cognition data. $6.65 per rollout, flag rate 1.3%
GPT-5.6 SolOpenAI47.5% (pass rate 52.9%)Codex, max, 5 trialsBoard 2026-10-06Cognition data. $5.19 per rollout, flag rate 0.0%
Claude Opus 4.8Anthropic46.5% (pass rate 51.6%)Claude Code, max, 5 trialsBoard 2026-10-06Cognition data. $9.62 per rollout, flag rate 0.6%
Kimi K3Moonshot44.2% (pass rate 48.9%)mini-swe-agent, no effort setting, 5 trialsAdded 2026-07-27Cognition data. $3.82 per rollout, flag rate 0.2%
CognitionMain, 100 tasks, v1.1cognition.com

"Added" dates come from Cognition's changelog and mark when a model joined the board, not its release date. "Board 2026-10-06" means the row was on the board when accessed and Cognition's changelog does not date its addition.

The top five sit within 2.5 points of each other, and four of them are Anthropic models in Claude Code. Below the top 14, the board continues with Gemini 3.7 Flash at 43.6%, GPT-5.5 at 43.0%, Claude Sonnet 5 at 42.7%, Grok 4.5 and GPT-6 Luna at 42.4%, SWE-1.7 at 42.0%, GPT-5.6 Terra at 41.3%, Gemini 3.8 Flash at 41.2% and GLM 5.3 at 40.1%. Mistral 3.5 Medium is last of 42 at 8.0%.

The two sources disagree on one point. Cognition's data has Fable 5.1 at 49.8% at Low effort ($2.38) and 50.9% at Medium, its best. Anthropic's Opus 5.5 launch page plots Fable 5.1 Low at 52.8% ($2.47). Medium through Max match between the two. On Anthropic's number Fable 5.1 would move above Sonnet 5.5 but stay below Opus 5. This page follows Cognition's file.

Extended, 150 tasks, best effort per model

Comparable to: other rows in this table (same version and subset, maintainer-run, harness and effort differ per row). Not comparable to Main, because Extended adds 50 easier tasks.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic65.3% (pass rate 70.8%)Claude Code, medium, 5 trials2026-10-06Cognition data. $0.67 per rollout
Claude Fable 5Anthropic64.9% (pass rate 70.9%)Claude Code, xhigh, 5 trials2026-10-06Cognition data. $10.53 per rollout
GPT-6 AstraOpenAI64.5% (pass rate 70.6%)Codex, max, 5 trials2026-10-06Cognition data. $3.93 per rollout
Claude Sonnet 5.5Anthropic64.4% (pass rate 69.9%)Claude Code, xhigh, 5 trials2026-10-06Cognition data. $1.24 per rollout
Claude Opus 5Anthropic63.6% (pass rate 69.6%)Claude Code, medium, 5 trials2026-10-06Cognition data. $3.51 per rollout
Claude Fable 5.1Anthropic63.6% (pass rate 68.8%)Claude Code, medium, 5 trials2026-10-06Cognition data. $2.68 per rollout
SWE-2Cognition62.5% (pass rate 68.4%)Devin, max, 5 trials2026-10-06Cognition data. $0.94 per rollout
Grok 4.6xAI61.3% (pass rate 67.0%)grok-build, high, 5 trials2026-10-06Cognition data. $2.38 per rollout
GPT-6 SolOpenAI60.7% (pass rate 66.3%)Codex, max, 5 trials2026-10-06Cognition data. $1.66 per rollout
GPT-5.6 SolOpenAI60.6% (pass rate 66.6%)Codex, max, 5 trials2026-10-06Cognition data. $4.25 per rollout
GPT-6.1 SolOpenAI60.4% (pass rate 66.5%)Codex, medium, 5 trials2026-10-06Cognition data. $0.31 per rollout
Claude Opus 4.8Anthropic59.6% (pass rate 65.5%)Claude Code, max, 5 trials2026-10-06Cognition data. $8.05 per rollout
Grok 4.7xAI59.4% (pass rate 65.2%)grok-build, high, 5 trials2026-10-06Cognition data. $5.59 per rollout
Kimi K3Moonshot58.2% (pass rate 63.6%)mini-swe-agent, no effort setting, 5 trials2026-10-06Cognition data. $3.12 per rollout
CognitionExtended, 150 tasks, v1.1cognition.com

Best effort is picked per subset, so a model's best Extended effort could differ from its best Main effort. For these 14 rows it doesn't. Each of the top ten Main models scores 10.2 to 13.3 points higher on Extended. Across those ten, the spread is 6.6 points on Main (54.6 to 48.0) and 4.9 points on Extended (65.3 to 60.4). Main separates the leaders a little better, which is why it is the subset worth quoting.

GPT-6.1 Sol drops from 7th on Main to 11th on Extended, while GPT-6 Astra climbs from 4th to 3rd. Neither move changes the cost story. GPT-6.1 Sol is still the cheapest of the top 14 on both subsets.

Devin Fusion runs on Extended

Comparable to: other rows in this table only. These runs use Cognition's Devin Fusion architecture, where a lead model delegates to a cheaper "sidekick" subagent, not Claude Code, so they are not comparable to the leaderboard tables above.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Fable 5 (low) + SidekickAnthropic60.7Devin Fusion lead with sidekick2026-07-13Cognition post, 3,000 sessions. $1.86 per run
Claude Fable 5 (low) aloneAnthropic60.8Devin Fusion lead, no sidekick2026-07-13Same post. $4.03 per run
Claude Opus 4.8 (medium) aloneAnthropic55.4Devin Fusion lead, no sidekick2026-07-13Same post. $3.06 per run
Claude Opus 4.8 (medium) + SidekickAnthropic54.6Devin Fusion lead with sidekick2026-07-13Same post. $2.04 per run
CognitionDevin FusionExtendedcognition.com

The sidekick cuts Fable 5's cost by more than half for a 0.1-point loss. Cognition explains the gap between the two leads by behavior. Fable 5 as lead takes 11.5 turns per run against 26.5 for Opus 4.8, and it makes no code edit itself in 81% of runs, against 24% for Opus 4.8. Fable 5 delegates. Opus 4.8 does the work itself and pays for it.

Effort settings change the ranking

The leaderboard picks each model's best effort, and for Claude models the best is rarely the maximum. The table below shows the full sweep for ten models on Main.

Comparable to: every cell in this table (same version, subset and harness per model). Not comparable to Extended.

ModelLowMediumHighXhighMax
Claude Opus 5.547.3 / $0.40 / 4.9 min54.6 / $0.80 / 6.3 min54.0 / $1.09 / 7.2 min51.4 / $2.25 / 11.2 min54.4 / $6.19 / 25.0 min
Claude Sonnet 5.529.3 / $0.19 / 3.8 min36.5 / $0.24 / 4.4 min49.4 / $0.42 / 5.9 min52.1 / $1.59 / 12.6 min46.2 / $20.78 / 62.4 min
GPT-6.1 Sol45.5 / $0.25 / 11.2 min50.2 / $0.36 / 14.4 min48.0 / $0.50 / 16.2 min49.3 / $0.56 / 17.1 min47.6 / $0.85 / 21.8 min
GPT-6 Astra45.3 / $1.7048.8 / $2.4350.9 / $3.0150.6 / $3.2853.3 / $4.59
Claude Fable 5.149.8 / $2.3850.9 / $3.2850.3 / $5.2748.7 / $9.2750.3 / $12.83
Claude Opus 541.9 / $2.6853.4 / $4.3148.0 / $7.2443.6 / $9.1448.0 / $11.42
Claude Fable 548.0 / $4.9249.8 / $7.1552.7 / $9.4853.5 / $13.0951.6 / $19.07
GPT-6 Sol37.3 / $0.4345.9 / $0.7747.7 / $1.0448.4 / $1.3349.3 / $2.07
SWE-2n/a43.1 / $0.3747.2 / $0.78n/a50.0 / $1.18
GPT-6 Luna25.7 / $0.0235.5 / $0.0537.3 / $0.0637.1 / $0.0742.4 / $0.10
CognitionMain, v1.1cognition.com

Cells are score / cost per rollout / minutes per rollout where Cognition publishes time. Source is Cognition's data file. Anthropic's Opus 5.5 page and Sonnet 5.5 page reproduce the Opus 5.5 and Sonnet 5.5 rows to the cent. Anthropic's Opus 5 medium and Astra max costs differ slightly ($4.61 and $4.36) because Cognition corrected Astra pricing on September 10.

Opus 5.5 peaks at medium, gives back 3.2 points at xhigh, and recovers to 54.4% at max for almost eight times the medium cost. Opus 5 falls nearly 10 points from medium to xhigh. Fable 5.1 peaks at medium. Sonnet 5.5 peaks at xhigh and falls to 46.2% at max, where it costs $20.78 and takes 62.4 minutes per task. GPT-6 Astra climbs with effort and peaks at max, as do GPT-6 Sol, GPT-6 Luna and SWE-2. More test-time compute buys nothing on this benchmark once a Claude model is past medium or xhigh.

Watch for this when reading launch pages. Anthropic's launch tables headline the Max column, so they show Opus 5.5 at 54.4, Opus 5 at 48.0, Fable 5.1 at 50.3, Sonnet 5.5 at 46.2, Astra at 53.3 and GPT-5.6 Sol at 47.5. Those are max-effort scores. Anthropic's own text gives Opus 5.5 at its default medium effort as 54.6%.

History: FrontierCode 1.0

Comparable to: other rows in the same column of this table. Not comparable to any 1.1 score, because the 1.1 blocker audit and internet verifier changed absolute scores.

ModelOrganizationDiamond (50)Main (100)Extended (150)HarnessSource and notes
Claude Fable 5Anthropicnot in post46.8%61.6%Claude CodeCognition data, v1 block
Claude Opus 4.8Anthropic13.4%34.3%51.8%Claude Code1.0 post and Cognition data
GPT-5.5OpenAI6.3%25.5%44.8%Codex1.0 post and Cognition data
Claude Opus 4.7Anthropicnot in post23.8%43.9%Claude CodeCognition data, v1 block
Kimi K2.7Moonshotnot in post22.0%39.9%mini-swe-agentCognition data, v1 block
Kimi K2.6Moonshot3.8%16.0%37.0%mini-swe-agent1.0 post, best open model in the post
Gemini 3.1 ProGoogle4.7%16.7%34.2%gemini-cli1.0 post and Cognition data
Cognition1.0: Diamond (50), Main (100), Extended (150)Main (100)

All rows are best effort, published June 8, 2026, before the fair-use prompt and the blocker audit. In the 1.0 post Cognition said Diamond "remains unsaturated," that Opus 4.8 held a clear lead, and that GPT-5.5 used up to four times fewer tokens than Opus 4.8.

Inside 1.1, where scores are comparable, the model lines climb steadily. On Main and Extended, the Opus line runs Opus 4.6 at 26.6 / 43.7, Opus 4.7 at 38.5 / 53.9, Opus 4.8 at 46.5 / 59.6, Opus 5 at 53.4 / 63.6 and Opus 5.5 at 54.6 / 65.3. The Sonnet line runs Sonnet 4.6 at 24.3 / 40.0, Sonnet 5 at 42.7 / 56.2 and Sonnet 5.5 at 52.1 / 64.4. The jump from Opus 4.8 to Opus 5 was 6.9 Main points. Opus 5 to Opus 5.5 was 1.2. Sonnet 5.5 gained 9.4 points over Sonnet 5 and now sits 2.5 points below Opus 5.5 on Main.

Failure modes and limitations

Contamination through the open internet

Every task comes from real PRs, so the fix exists somewhere in a later upstream release, a mirror or a package registry. An agent can install the latest version of a dependency and find the answer sitting in it. Cognition saw rare cases in 1.0 and a rising rate in newer models, naming Fable 5, which is why 1.1 added the fair-use prompt and verifier. Flag rates in the data run from 0.0% for Opus 5.5 on Main to 25.5% for DeepSeek V4 Flash 0731 and 10.6% for DeepSeek V4 Pro 0813. A flagged run scores 0, so a model with a high flag rate is losing points to the rule, not only to bad patches.

Over-scoped edits at high effort

Anthropic's footnote on the Sonnet 5.5 page explains the Max drop directly: "Sonnet 5.5 scores lower at Max effort than at Xhigh. FrontierCode evaluates whether a code change could be merged without human edits. It penalizes out-of-scope changes, even if they are high-quality or helpful. At Max effort, Sonnet 5.5 more often ran Claude Code's code-review skill, which splits the review across many subagents, and in two cases Cognition examined, this led to a timeout or to extra edits beyond the task's scope, and therefore to a lower score."

That is the scope axis doing its job. A model that does more than it was asked fails the review a maintainer would give it. Teams that run agents at maximum effort by default should take the hint.

Subjective rubrics

Cognition says plainly that rubric grading is subjective and that its QC process is the defense. The 1.1 audit found 75 of more than 1,000 blocker criteria too strict. That is a correction rate of up to 7.5%, found by the people who wrote the rubrics. Prompt graders add an LLM judge to the loop, with its own thresholds.

Harness confound

Claude models run in Claude Code and GPT models in Codex. Cognition publishes no FrontierCode result for any other model and agent pairing, so the board cannot separate a model's ability from its vendor's agent.

No outside reproduction

Tasks, rubrics and harness configs are private, and evaluation is open to model creators only. Cognition publishes five-trial averages with no confidence intervals or run-to-run variance. With 100 tasks in Main and five trials each, a gap like Opus 5.5's 1.1-point lead over Fable 5 comes with no published error bar. And Cognition grades its own SWE-2, which places 8th on Main.

Small task count and a dropped subset

The whole benchmark is 150 tasks. Diamond was dropped because its solve rates were too low to give a stable signal. Main, at 100 tasks, is the hardest subset left.

FrontierCode belongs with the agentic coding benchmarks, and it differs from them mainly in what counts as success. For a broader map, see LLM benchmarks and LLM evaluation.

BenchmarkWhat it scoresHow the leaders compare
FrontierCode 1.1 MainWhether a maintainer would merge the patch, six-axis rubricOpus 5.5 54.6% (medium) vs GPT-6 Astra 53.3% (max), 1.3 points
SWE-bench Verified and ProHidden tests pass or fail per issueCognition reports an 81% lower false positive rate for FrontierCode than SWE-Bench Pro
Terminal-Bench 4.0Resolution rate on terminal tasks, with cost and token columnsAnthropic reports Opus 5.5 66.4% (xhigh, ±2.6) vs GPT-6 Astra 57.9% (high), 8.5 points
CursorBenchIn-editor sessions traced with Cursor Blame, agentic gradersAnthropic reports CursorBench 4.0 Opus 5.5 57.8% (max), Fable 5.1 51.8%, Opus 5 46.6%, no GPT-6 Astra score
DeepSWELarger patches than FrontierCodeCognition says FrontierCode is harder despite smaller patches

Cognition built FrontierCode in response to SWE-bench's limits. SWE-bench's pass/fail tests miss the scope and code quality problems that get PRs rejected. FrontierCode grades those directly, and its adaptive classical graders accept alternative valid implementations that SWE-bench's fixed tests would reject.

Terminal-Bench 4.0, run by Stanford, Harbor and the Laude Institute, measures shell competence. The same two models are 8.5 points apart there and 1.3 apart on FrontierCode Main. Opus 5.5's lead holds on both, but a team choosing between Opus 5.5 and Astra for patch work gets much less separation from FrontierCode than the Terminal-Bench gap suggests.

CursorBench is Cursor's internal suite, built from real sessions with ambiguous prompts and refreshed every few months, with scores comparable only within a version. Cursor's post names version 3.1, and Anthropic reports 4.0. It tests the interactive editor workflow. FrontierCode tests the autonomous path from issue to patch, graded by maintainer rubrics.

FrontierSWE v2 and ProgramBench are the other sibling pages in this set. FrontierSWE runs multi-hour open-ended projects, a different task length from FrontierCode's single mergeable patch. For the older function-level benchmarks FrontierCode moves away from, see HumanEval and BigCodeBench.

What it means for teams choosing a model

Opus 5.5 at medium effort in Claude Code is the default pick for mergeable patch work. It leads Main at 54.6% and Extended at 65.3%, costs $0.80 per task on Main and finishes in 6.3 minutes. GPT-6 Astra at max effort in Codex is 1.3 points behind on Main and 0.8 behind on Extended, at $4.59 per task. Unless your team is committed to Codex, Astra buys nothing here that Opus 5.5 doesn't deliver for less.

If cost matters more than the last few points, GPT-6.1 Sol at medium is the budget choice. It scores 50.2% for $0.36, 4.4 points under Opus 5.5 at 45% of its cost, though it takes 14.4 minutes per task against 6.3. GPT-6 Luna at max scores 42.4% for $0.10, which is a reasonable floor for high-volume, low-stakes changes. Sonnet 5.5 at high (49.4%, $0.42, 5.9 minutes) is the fast alternative in Claude Code.

Don't set effort to max because it sounds safest. Every Claude model in the effort sweep peaks below max. Sonnet 5.5 at max spends 13 times the money and five times the minutes of its xhigh setting to score 5.9 points lower, because it does more than the task asked. Astra is the exception, and its best score needs max.

Fable 5 at xhigh scores 53.5% for $13.09, more than 16 times Opus 5.5's medium cost for a lower score. Cognition's own Devin Fusion runs show where Fable 5 earns its keep, as a lead that delegates to a cheaper subagent, at 60.7% on Extended for $1.86.

Treat each score as a model plus an agent. Every Claude number comes from Claude Code and every GPT number from Codex. If you plan to run a model in a different agent, FrontierCode has no number for that setup. And the scores come from Cognition alone, with no error bars, so a one-point gap at the top is a tie for practical purposes.

The frontier models are close on FrontierCode, and your repository's conventions are not the same as 36 open-source projects'. You can compare current models on the LLM leaderboard and run the same kind of rubric-graded, scope-aware evaluation on your own repositories in Klu, as described on the engineering page.