CursorBench

Cursor's private coding-agent benchmark, built from real Cursor sessions, that scores ambiguous multi-file tasks on correctness, code quality, efficiency, and interaction behavior

Agentic coding

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 12 min read

What is CursorBench?

CursorBench is the benchmark Cursor (Anysphere) built to compare model quality inside its editor. It scores coding agents on ambiguous, multi-file tasks taken from real Cursor sessions. Cursor launched it as CursorBench-3 on 2026-03-11 in the post How we compare model quality in Cursor and moved the suite to version 4.0 on 2026-09-10. The public leaderboard scores 63 model-and-effort configurations on the 4.0 set and prints the average cost per task beside every score.

That cost column is why the benchmark is worth reading. Most LLM benchmarks give you an accuracy number and leave the bill as an exercise. CursorBench gives you both from one run, so you can see that Opus 5.5 at High effort scores 56.0% for $3.97 per task while the same model at Max scores 57.8% for $13.43.

On the 4.0 table, Opus 5.5 Max leads at 57.8%. Anthropic holds all of the top 12 rows. The best configuration from any other lab is Grok 4.7 Extra High at 46.3%, 11.5 points back, and OpenAI's best is GPT-5.6 Sol Max at 41.7%, 16.1 points back.

The "3.2.0" in this page's title is the previous task set. Cursor's leaderboard now shows only 4.0, and the only 3.2 score with a primary source is a launch-partner quote in Anthropic's Fable 5.1 announcement. The two sets are kept in separate tables below.

Cursor's launch post names three problems with public benchmarks that CursorBench is meant to fix. Public tasks don't match the work developers do. They are hard to grade, because "many public benchmark tasks assume a narrow set of correct solutions, but most developer requests are underspecified enough to admit many valid approaches." And they leak into training data. Cursor points to OpenAI having stopped reporting SWE-bench Verified results over contamination.

How the tasks are built

Cursor sources tasks with a tool it calls Cursor Blame, which "traces committed code back to the agent request that produced it." That gives Cursor a natural pair for every task: the request a developer typed, and the code that request eventually produced and that the developer committed. Nobody writes a synthetic problem statement.

The prompts are "intentionally short, in contrast with the detailed GitHub issues sourced in public benchmarks." This is the design choice with the most consequence. A SWE-bench task hands the agent an issue written for a human maintainer. A CursorBench task hands it the kind of terse request a developer types into an agent pane, and the agent has to work out what was meant before it can do anything. Tasks involve "multiple files, tools, and steps." Cursor does not list the tools.

Task scope has grown with each version. Cursor says scope "roughly doubled" in lines changed and mean files touched between the first version and CursorBench-3, and that CursorBench-3 tasks involve "substantially more lines than those in SWE-bench Verified, Pro, or Multilingual." Version 4.0 added long-horizon problems on top of that.

Contamination is handled by design rather than by audit. Many tasks come from Cursor's internal codebases and other controlled sources "to reduce training data contamination risk," and Cursor refreshes the suite every few months. The tasks are private, which keeps them out of training sets and also means no one outside Cursor can re-run them.

How scoring works

"Agentic graders" score each task on four dimensions: solution correctness, code quality, efficiency and interaction behavior. This is a form of LLM-as-a-judge evaluation. Cursor uses it because its short requests "admit many valid approaches," and grading against "a narrow set of correct solutions" is the problem it set out to fix in public benchmarks. Cursor has not published the grader model, the rubric weights or the per-task pass criteria, so the 57.8% at the top of the table cannot be split into how much came from correctness and how much from code quality.

Cost is computed, not billed. Cursor's method, quoted from the leaderboard: "Avg cost / task is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task." Cost, tokens and steps are all per-task averages. Cache writes count toward the total, and Cursor updated GPT-5.6 Sol, Terra and Luna on 2026-07-09 for cache write costs.

Reasoning effort is its own column, with levels Minimal, Low, Medium, High, Extra High and Max. Not every model has every level. This is where test-time compute shows up in the table: the same model at different effort levels appears as separate rows. Opus 5.5 spans 14.1 points from Low (43.7%) to Max (57.8%), more than the 11.5 points between the top Anthropic row and the best row from any other lab. Cursor has not published the token budget behind each level.

CursorBench is half of Cursor's evaluation process. The other half is online controlled experiments on live traffic. Cursor says online evals "catch regressions that offline suites miss, like where the agent's output looks correct to a grader but feels worse to a developer using the product." In one example, Cursor removed semantic search entirely to measure where it helps most on repository question-answering in larger codebases. The leaderboard publishes only the offline half.

Cursor has not published the task count, the number of tasks per category, the tool list, time or step limits, the sandbox spec, the number of runs per configuration, confidence intervals, or whether the benchmark harness matches the one Cursor uses in production or for RL. Its only statement on noise is one line: "Results are subject to variance; small differences in scores may not be statistically meaningful."

Versions

Cursor's changelog, identical on cursor.com/evals and cursor.com/cursorbench, lists four task-set releases.

VersionDateWhat Cursor added
3.02026-03-11"Initial set of tasks focused on edit, refactor, and bugfix problems"
3.12026-05-19"problems focused on codebase understanding, bugfinding, planning, and code review"
3.22026-07-08"instruction following and advanced tool use problems"
4.02026-09-10"new long-horizon problems focused on edit, refactor, investigation, intent understanding, managing jobs, and design adherence"

The same pages list three cost corrections made after publication: GPT-5.6 Sol, Terra and Luna on 2026-07-09 for cache write costs, GPT-5.6 Terra and Luna on 2026-07-30 for adjusted pricing, and Sonnet 5 on 2026-08-11 for adjusted pricing. If you saved a copy of the table before 2026-08-11, its cost column for those models is out of date.

Each release added categories, so every version is a different test. Cursor makes no statement on whether 4.0 scores compare with 3.x scores, and its leaderboard no longer shows a 3.x table. A 3.2 number and a 4.0 number for the same model tell you nothing about whether that model got better or worse.

Current leaderboard

CursorBench 4.0, Cursor's run (63 configurations)

Run and published by Cursor, independent of the model vendors, on the 4.0 task set; comparable only within this table. Cursor does not publish its harness, tool list or limits, gives no per-row evaluation date, and does not name organizations, so the Organization column is assigned by model family. The table includes Opus 5.5, released 2026-09-22 per Anthropic. No GPT-6 Astra, GPT-6.1, Gemini 4 or Claude Mythos configuration appears.

ModelOrganizationScoreHarness and setupDateSource and notes
Opus 5.5Anthropic57.8%Cursor's run, Max effortAccessed 2026-10-06cursor.com/cursorbench. Rank 1. $13.43/task, 218,363 tokens, 185 steps
Opus 5.5Anthropic56.0%Cursor's run, Extra High effortAccessed 2026-10-06Rank 2. $6.98/task, 101,083 tokens, 109 steps
Opus 5.5Anthropic56.0%Cursor's run, High effortAccessed 2026-10-06Rank 3. $3.97/task, 53,078 tokens, 68 steps
Sonnet 5.5Anthropic55.5%Cursor's run, Max effortAccessed 2026-10-06Rank 4. $9.67/task, 271,920 tokens, 170 steps
Sonnet 5.5Anthropic53.1%Cursor's run, Extra High effortAccessed 2026-10-06Rank 5. $3.88/task, 100,158 tokens, 78 steps
Opus 5.5Anthropic52.5%Cursor's run, Medium effortAccessed 2026-10-06Rank 6. $2.91/task, 37,954 tokens, 54 steps
Fable 5.1Anthropic51.8%Cursor's run, Max effortAccessed 2026-10-06Rank 7. $17.28/task, 117,236 tokens, 128 steps
Fable 5.1Anthropic51.6%Cursor's run, Extra High effortAccessed 2026-10-06Rank 8. $13.01/task, 87,294 tokens, 101 steps
Fable 5.1Anthropic49.2%Cursor's run, High effortAccessed 2026-10-06Rank 9. $9.08/task, 58,438 tokens, 77 steps
Sonnet 5.5Anthropic47.8%Cursor's run, High effortAccessed 2026-10-06Rank 10. $1.67/task, 37,391 tokens, 41 steps
Fable 5.1Anthropic46.8%Cursor's run, Medium effortAccessed 2026-10-06Rank 11. $7.05/task, 45,411 tokens, 63 steps
Opus 5Anthropic46.6%Cursor's run, Max effortAccessed 2026-10-06Rank 12. $11.95/task, 85,384 tokens, 106 steps
Grok 4.7xAI46.3%Cursor's run, Extra High effortAccessed 2026-10-06Rank 13. $6.01/task, 70,141 tokens, 88 steps
Opus 5Anthropic46.1%Cursor's run, Extra High effortAccessed 2026-10-06Rank 14. $11.43/task, 80,094 tokens, 103 steps
Fable 5.1Anthropic45.1%Cursor's run, Low effortAccessed 2026-10-06Rank 15. $5.44/task, 34,795 tokens, 51 steps
Opus 5Anthropic44.7%Cursor's run, High effortAccessed 2026-10-06Rank 16. $9.00/task, 61,405 tokens, 86 steps
Grok 4.7xAI43.9%Cursor's run, High effortAccessed 2026-10-06Rank 17. $4.69/task, 56,382 tokens, 71 steps
Opus 5.5Anthropic43.7%Cursor's run, Low effortAccessed 2026-10-06Rank 18. $1.17/task, 15,811 tokens, 28 steps
Opus 5Anthropic43.3%Cursor's run, Medium effortAccessed 2026-10-06Rank 19. $6.94/task, 45,272 tokens, 72 steps
GLM 5.3Z.ai42.6%Cursor's run, Max effortAccessed 2026-10-06Rank 20. $5.05/task, 96,387 tokens, 166 steps
GPT-5.6 SolOpenAI41.7%Cursor's run, Max effortAccessed 2026-10-06Rank 21. $8.23/task, 42,944 tokens, 99 steps
Grok 4.7xAI41.6%Cursor's run, Medium effortAccessed 2026-10-06Rank 22. $3.49/task, 36,683 tokens, 60 steps
Muse Spark 1.3Not stated41.6%Cursor's run, Max effortAccessed 2026-10-06Rank 23. $2.64/task, 52,005 tokens, 98 steps
Grok 4.6xAI41.4%Cursor's run, Extra High effortAccessed 2026-10-06Rank 24. $6.10/task, 49,814 tokens, 56 steps
GPT-5.6 TerraOpenAI41.3%Cursor's run, Max effortAccessed 2026-10-06Rank 25. $5.14/task, 60,814 tokens, 107 steps
Opus 5Anthropic40.7%Cursor's run, Low effortAccessed 2026-10-06Rank 26. $4.87/task, 31,995 tokens, 57 steps
Grok 4.6xAI40.4%Cursor's run, High effortAccessed 2026-10-06Rank 27. $5.20/task, 41,387 tokens, 48 steps
Gemini 3.8 FlashGoogle39.6%Cursor's run, High effortAccessed 2026-10-06Rank 28. $4.70/task, 162,565 tokens, 324 steps
Sonnet 5.5Anthropic39.2%Cursor's run, Medium effortAccessed 2026-10-06Rank 29. $0.70/task, 16,036 tokens, 22 steps
GLM 5.3Z.ai38.0%Cursor's run, High effortAccessed 2026-10-06Rank 30. $3.24/task, 60,031 tokens, 114 steps
GPT-5.6 SolOpenAI37.7%Cursor's run, Extra High effortAccessed 2026-10-06Rank 31. $4.40/task, 24,729 tokens, 55 steps
Muse Spark 1.3Not stated37.5%Cursor's run, Extra High effortAccessed 2026-10-06Rank 32. $2.10/task, 40,891 tokens, 83 steps
Gemini 3.8 FlashGoogle37.3%Cursor's run, Medium effortAccessed 2026-10-06Rank 33. $4.06/task, 128,364 tokens, 290 steps
GLM 5.3 FlashZ.ai36.8%Cursor's run, Max effortAccessed 2026-10-06Rank 34. $0.39/task, 56,410 tokens, 118 steps
Grok 4.6xAI36.1%Cursor's run, Medium effortAccessed 2026-10-06Rank 35. $3.48/task, 24,893 tokens, 40 steps
GPT-5.6 LunaOpenAI35.9%Cursor's run, Max effortAccessed 2026-10-06Rank 36. $1.03/task, 87,284 tokens, 208 steps
Sonnet 5.5Anthropic35.8%Cursor's run, Low effortAccessed 2026-10-06Rank 37. $0.50/task, 11,668 tokens, 18 steps
GPT-5.6 SolOpenAI35.7%Cursor's run, High effortAccessed 2026-10-06Rank 38. $2.85/task, 16,174 tokens, 41 steps
Sonnet 5Anthropic34.1%Cursor's run, Max effortAccessed 2026-10-06Rank 39. $7.17/task, 149,257 tokens, 140 steps
GPT-5.6 TerraOpenAI33.6%Cursor's run, Extra High effortAccessed 2026-10-06Rank 40. $1.81/task, 23,436 tokens, 43 steps
Grok 4.6xAI33.4%Cursor's run, Low effortAccessed 2026-10-06Rank 41. $2.25/task, 16,307 tokens, 32 steps
Muse Spark 1.3Not stated33.4%Cursor's run, High effortAccessed 2026-10-06Rank 42. $1.66/task, 30,654 tokens, 69 steps
GLM 5.3Z.ai33.3%Cursor's run, Low effortAccessed 2026-10-06Rank 43. $2.04/task, 31,983 tokens, 81 steps
Grok 4.7xAI33.1%Cursor's run, Low effortAccessed 2026-10-06Rank 44. $1.58/task, 15,677 tokens, 40 steps
GPT-5.6 LunaOpenAI33.0%Cursor's run, Extra High effortAccessed 2026-10-06Rank 45. $0.44/task, 40,598 tokens, 98 steps
Muse Spark 1.3Not stated32.6%Cursor's run, Medium effortAccessed 2026-10-06Rank 46. $1.49/task, 27,255 tokens, 64 steps
Sonnet 5Anthropic32.0%Cursor's run, Extra High effortAccessed 2026-10-06Rank 47. $4.55/task, 83,373 tokens, 102 steps
GPT-5.6 SolOpenAI31.1%Cursor's run, Medium effortAccessed 2026-10-06Rank 48. $1.77/task, 10,111 tokens, 32 steps
GLM 5.3 FlashZ.ai31.1%Cursor's run, High effortAccessed 2026-10-06Rank 49. $0.25/task, 35,104 tokens, 84 steps
Sonnet 5Anthropic30.8%Cursor's run, High effortAccessed 2026-10-06Rank 50. $3.48/task, 61,146 tokens, 85 steps
GPT-5.6 TerraOpenAI30.7%Cursor's run, High effortAccessed 2026-10-06Rank 51. $1.11/task, 13,162 tokens, 33 steps
GPT-5.6 LunaOpenAI29.4%Cursor's run, High effortAccessed 2026-10-06Rank 52. $0.25/task, 23,368 tokens, 64 steps
Muse Spark 1.3Not stated29.3%Cursor's run, Low effortAccessed 2026-10-06Rank 53. $0.93/task, 17,483 tokens, 47 steps
Sonnet 5Anthropic28.0%Cursor's run, Medium effortAccessed 2026-10-06Rank 54. $2.31/task, 39,114 tokens, 65 steps
Composer 2.5Cursor27.7%Cursor's run, no effort level listedAccessed 2026-10-06Rank 55. $0.68/task, 17,347 tokens, 41 steps
GPT-5.6 TerraOpenAI27.6%Cursor's run, Medium effortAccessed 2026-10-06Rank 56. $0.64/task, 7,307 tokens, 25 steps
GLM 5.3 FlashZ.ai26.9%Cursor's run, Low effortAccessed 2026-10-06Rank 57. $0.15/task, 17,831 tokens, 58 steps
GPT-5.6 TerraOpenAI25.2%Cursor's run, Low effortAccessed 2026-10-06Rank 58. $0.52/task, 5,914 tokens, 23 steps
GPT-5.6 SolOpenAI24.6%Cursor's run, Low effortAccessed 2026-10-06Rank 59. $0.87/task, 4,885 tokens, 21 steps
Muse Spark 1.3Not stated24.3%Cursor's run, Minimal effortAccessed 2026-10-06Rank 60. $0.56/task, 10,620 tokens, 34 steps
Sonnet 5Anthropic24.1%Cursor's run, Low effortAccessed 2026-10-06Rank 61. $1.39/task, 23,772 tokens, 46 steps
GPT-5.6 LunaOpenAI22.2%Cursor's run, Medium effortAccessed 2026-10-06Rank 62. $0.08/task, 7,642 tokens, 32 steps
GPT-5.6 LunaOpenAI16.0%Cursor's run, Low effortAccessed 2026-10-06Rank 63. $0.03/task, 3,288 tokens, 18 steps
Cursor4.0cursor.com
CursorBench 4.0: score vs cost per taskCursor's own run on the 4.0 task set, accessed 2026-10-06
40%50%60%$0.5$1$2$5$10$20
FrontierOtherAnthropicBehind the frontier
Cost per task, log scale, lower to the right
Source: cursor.com/cursorbench, accessed 2026-10-06. Cursor runs every configuration and computes cost per task by applying each model's published per-million-token prices (input, cache read, cache write, output) to the tokens used; no further derivation here. Harness not published.

Opus 5.5 leads at every effort level from Medium up. Its Max configuration tops the table at 57.8%, and Sonnet 5.5 Max is next among other models at 55.5%, 2.3 points behind. Cursor warns that small differences "may not be statistically meaningful" and publishes no confidence intervals, so that gap is unresolved. What is not in doubt is cost: Sonnet 5.5 Max takes 271,920 tokens per task, the most of any configuration, and costs $9.67 to Opus 5.5 High's $3.97 while scoring 0.5 points lower.

The cost frontier, the set of configurations that no other row beats on score at equal or lower cost, belongs to Anthropic's 5.5 family from $0.70 upward. In order: Sonnet 5.5 Medium (39.2%, $0.70), Opus 5.5 Low (43.7%, $1.17), Sonnet 5.5 High (47.8%, $1.67), Opus 5.5 Medium (52.5%, $2.91), Sonnet 5.5 Extra High (53.1%, $3.88), Opus 5.5 High (56.0%, $3.97) and Opus 5.5 Max (57.8%, $13.43). Below $0.70 the frontier is GLM 5.3 Flash and GPT-5.6 Luna, with GLM 5.3 Flash Max at 36.8% for $0.39 the strongest cheap option.

Opus 5.5 Extra High is the oddest row. It ties Opus 5.5 High at 56.0% and costs 1.76x as much, $6.98 against $3.97. Going from High to Extra High buys nothing on this set. Going from High to Max buys 1.8 points for 3.4x the cost.

Fable 5.1 is off the frontier at every effort level. Its best score is 51.8% at Max for $17.28, the most expensive configuration on the board, and Opus 5.5 Medium beats it at 52.5% for $2.91. The Fable price gap is per token, not per task length. Fable 5.1 Max uses 117,236 tokens per task against Opus 5.5 Max's 218,363 and still costs $3.85 more.

OpenAI's GPT-5.6 Sol runs lean and scores low here. Sol Max uses 42,944 tokens per task, a fifth of Opus 5.5 Max's count, and scores 41.7% at $8.23. Anthropic's Opus 5.5 announcement makes the same comparison: "It beats GPT-5.6 Sol's top score (41.7%) by 11 points for about a third of the cost per task," referring to Opus 5.5 at its default medium effort. Cursor's table agrees with Anthropic's figures.

Step counts vary more than scores. Gemini 3.8 Flash High takes 324 steps per task, the most of any configuration, to reach 39.6% at $4.70. Sonnet 5.5 Low reaches 35.8% in 18 steps. If your agent runs inside a product where each step is a visible action or a tool call with its own latency, the steps column matters as much as the cost column.

The table does not flatter its publisher. Cursor's own model, Composer 2.5, scores 27.7% at $0.68 per task, rank 55 of 63.

CursorBench 3.2, vendor-reported

Separate, retired task set; not comparable with the 4.0 table. This figure is a launch-partner quote in a vendor announcement with no harness stated, and it is the only 3.2 score with a primary source.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Fable 5.1Anthropic73.4%Max effort, harness not statedNot statedAnthropic's Fable 5.1 page. Vendor-reported partner quote: "Claude Fable 5.1 is the most capable model we've run on CursorBench 3.2, scoring 73.4% at max effort."
Anthropic3.2anthropic.com

Other 3.2 figures circulate on aggregator sites, including numbers for Fable 5, Grok 4.5, Grok 4.6, Opus 5 and Composer 2.5. None traces to Cursor or a lab, so this page leaves them out. Fable 5.1's 73.4% on 3.2 and its 51.8% on 4.0 come from different task sets and different reporters. Neither number says anything about the other.

Historical progression

Cursor publishes no earlier leaderboards, so the usable history is three Anthropic generations measured on the same 4.0 set in the same Cursor run.

ModelEffortScoreCost/taskSourceFrontier
Opus 5Max46.6%$11.95cursor.com/cursorbench
Fable 5.1Max51.8%$17.28Same table
Opus 5.5Max57.8%$13.43Same tableFrontier
Opus 5Medium43.3%$6.94Same table
Opus 5.5Medium52.5%$2.91Same tableFrontier
Cursor4.0cursor.com

At Max, Opus 5.5 leads Opus 5 by 11.2 points for $1.48 more per task. At Medium the move is larger. Opus 5.5 scores 9.2 points higher at 42% of Opus 5's cost. Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens, per Anthropic.

Documented failure modes and limitations

Small gaps are noise by Cursor's own account. "Small differences in scores may not be statistically meaningful," and Cursor publishes neither confidence intervals nor the number of runs per configuration. Opus 5.5 High and Extra High tie at 56.0%. The 2.3 points between Opus 5.5 Max and Sonnet 5.5 Max do not support a ranking. The 11.5 points between Opus 5.5 Max and Grok 4.7 Extra High do.

The grader is a black box. Cursor names "agentic graders" once and stops there. Without the grader model or rubric weights, you can't tell whether a model lost points on correctness or on code quality, and you can't check whether the grader favors the style of one model family. Cursor itself lists grading difficulty as a known problem and names reducing grading cost as future work.

Nobody outside Cursor can audit the scores. Private tasks are the contamination defense, and Cursor published no contamination analysis beyond that design. The same privacy means no lab, researcher or customer can re-run a configuration and check the number.

Offline scores miss things developers notice. Cursor says so directly, which is why it runs online experiments on live traffic alongside the benchmark. An agent can produce output that "looks correct to a grader but feels worse to a developer using the product." CursorBench measures the first half of that sentence.

External services hurt reproducibility. Cursor named reproducibility with external services as an open problem in its March post and has not said whether 4.0 addresses it.

The test and the prices move. Categories changed at every release from 3.0 to 4.0, and Cursor revised cost figures after publication on 2026-07-09, 2026-07-30 and 2026-08-11. Any cost comparison you make should cite the access date.

SWE-bench Verified, Pro and Multilingual draw tasks from public repositories and give the agent a detailed GitHub issue. CursorBench draws from private sessions and gives the agent a short request. Cursor says its tasks involve substantially more lines than all three SWE-bench variants, and it cites OpenAI's decision to stop reporting SWE-bench Verified over contamination as one reason to move to private tasks.

Terminal-Bench 4.0 is a public shell benchmark with a published task count: 66 new containerized terminal tasks, none overlapping the 89 in version 2.1, with an 8-hour agent timeout. Vals runs it with mini-swe-agent and reports avg@3, and lists Claude Opus 5.5 first at 65.15% and $13.20 per test as of 2026-10-01. Terminal-Bench is public shell work with a known task count. CursorBench is private editor-session work scored on code quality and interaction as well as correctness. Opus 5.5 leads both, on different tasks with different harnesses, so the two scores don't combine.

DeepSWE has 113 original tasks across 91 open-source repositories in five languages, written from scratch and never upstreamed, run with no internet in the sandbox. Hand-written behavioral verifiers grade it, and the DeepSWE paper reports a verifier-versus-LLM-judge disagreement of 1.4%, against 32.4% for SWE-Bench Pro. Its leaderboard has a three-way tie at 74% with overlapping confidence intervals: GPT-6 Astra xhigh at $4.43, Gemini 3.8 Flash high at $2.36 and Claude Opus 5 max at $11.84. DeepSWE publishes confidence intervals and program verifiers. CursorBench publishes neither. The rankings also disagree: Gemini 3.8 Flash High shares DeepSWE's top spot and sits at rank 28 on CursorBench 4.0. A team that read only one of the two would choose differently. GPT-6 Astra is not in Cursor's table at all.

FrontierCode, from Cognition, is the closest relative. It grades on six dimensions including regression safety and code quality, and it keeps its tasks private. Its 150 Extended, 100 Main and 50 Diamond tasks were built by more than 20 open-source maintainers at more than 40 hours per task. The difference is source material. FrontierCode tasks come from maintainers' repositories, CursorBench tasks from real editor sessions.

For the general background on how these evaluations fit together, see the LLM evaluation overview and the frontier models page.

What CursorBench means for teams choosing a model

Opus 5.5 High is the default pick for Cursor-style agent work. It scores 56.0% at $3.97 per task, 4.2 points above Fable 5.1 Max at 23% of Fable's $17.28. Only Opus 5.5 Max scores higher, and Sonnet 5.5 Max, the next-best model, costs more and scores lower.

Pay for Max only when failure is expensive. On Opus 5.5, Max adds 1.8 points over High for $9.46 more per task. That is 0.018 more probability of success, so Max pays for itself only where a failed task costs you more than about $525. For most interactive coding, it doesn't. For an unattended agent on a migration, where one failure costs an engineer a day of cleanup, it does.

Skip Extra High on Opus 5.5. It ties High at 56.0% and costs 1.76x as much.

Opus 5.5 Medium replaces Fable 5.1 outright. Medium is Anthropic's own default effort. It scores 52.5% at $2.91, beating every Fable 5.1 configuration, including Max, at 17% of Fable 5.1 Max's cost. On 4.0, every Fable 5.1 row is beaten by an Opus 5.5 row that costs less.

On a budget, stay in the Anthropic 5.5 family. Sonnet 5.5 High scores 47.8% at $1.67 and Opus 5.5 Low scores 43.7% at $1.17. Both beat GPT-5.6 Sol Max's 41.7%, and Opus 5.5 Low does it at 14% of Sol Max's $8.23.

For bulk, low-stakes work, look at GLM 5.3 Flash and GPT-5.6 Luna. GLM 5.3 Flash Max scores 36.8% at $0.39 and GPT-5.6 Luna Max scores 35.9% at $1.03. Gemini 3.8 Flash High scores higher at 39.6% but costs $4.70 and takes 324 steps per task, the most of any configuration, a poor fit wherever each step adds latency.

CursorBench is one signal. It measures Cursor's tasks, in Cursor's unpublished harness, scored by Cursor's unpublished graders. Cursor itself pairs it with live experiments because offline scores miss regressions users feel. Rankings shift across benchmarks too, as the Gemini 3.8 Flash result on DeepSWE shows.

The Klu LLM leaderboard tracks these models across other benchmarks. To run the same kind of cost-and-quality eval on your own codebase and prompts, see how engineering teams set it up in Klu on the engineering page.