What is BenchCAD?
BenchCAD tests whether a model can write parametric CAD code that reproduces a real mechanical part. The model sees the part and has to produce a CadQuery program. The program runs, the resulting solid is compared with the reference solid, and the score is how much of the two volumes overlap. No LLM judge is involved.
Haozhe Zhang, Kaichen Liu, Miaomiao Chen, Lei Li, Shaojie Yang, Cheng Peng and Hanjie Chen posted the paper to arXiv on 11 May 2026 as arXiv:2605.10865. The dataset is 17,900 execution-verified CadQuery programs across 106 industrial part families: gears, springs, twist drills, fasteners, structural parts, fluid fittings, panels and enclosures. 52 of those families bind their dimensions to ISO, DIN, EN, ASME or IEC specification tables, drawing on 47 distinct standard codes. The other 54 are custom parts.
The authors built it because earlier text-to-CAD and image-to-CAD sets are small, synthetic, or limited to sketch-and-extrude shapes. On those sets, the abstract says, models "often recover coarse outer geometry but fail to produce faithful parametric CAD programs." They miss fine 3D structure, misread design parameters, and replace sweeps, lofts and twist-extrudes with plain extrusions. BenchCAD covers 49 CadQuery operations, and its four tasks pull apart three skills that a single image-to-code score blends together: recognizing what is in the picture, abstracting it into parameters, and writing the code.
For frontier model comparisons, one task matters. Vision2Code gives the model four orthographic renders and asks for CadQuery code, scored by mean voxel IoU. Anthropic reports it in the Fable 5 and Mythos 5, Fable 5.1 and Mythos 5.1, Opus 5.5 and Sonnet 5.5 system cards, and those cards are the source of nearly every current frontier number.
How the tasks work
Every part in the dataset ships with its CadQuery code, a STEP file, four rendered views (front, top, right and isometric), a JSON file of design parameters and the list of operations the code uses. Each family has easy, medium and hard tiers. Two companion sets reuse the parts: BenchCAD-QA has 2,400 paired numeric questions and BenchCAD-Edit has 748 before-and-after edit pairs.
The four tasks are:
- Vision2Code (img2cq). Four renders in, a CadQuery program out. This is the headline task.
- Code Edit. The model gets a working program and an edit request, and must produce the edited program. Edits run from tier T1, a single literal swap, to T5, a multi-block restructure.
- Vision QA (qa_img). Numeric questions about a part answered from the images alone.
- Code QA (qa_code). The same questions answered from the CadQuery source.
The QA pair is the clever part. The questions match, so the gap between Vision QA and Code QA isolates how much a model loses by having to see the part instead of reading its definition. The paper grades QA by capability level, from L1 (holistic visual recognition) to L4 (spatial and code reasoning).
The execution sandbox gives each program 30 seconds. Outputs with degenerate geometry, a volume at or below 1e-6 cubic millimeters, are filtered out.
How scoring works
Vision2Code scoring voxelizes the predicted solid and the reference solid and divides the volume of their intersection by the volume of their union. A perfect reconstruction scores 1.0. Because the metric is continuous, a model that gets the outer envelope right and misses a small bore still earns most of the credit, so frontier models can post high IoU while dropping fine detail.
The paper reports more than IoU for Vision2Code. Its Table 10 adds Chamfer distance, a feature and essential-operation score, execution rate, and a combined total. Frontier labs report mean voxel IoU only, so the operation-level checks that would catch a twist-extrude replaced by a straight extrude never appear in the vendor numbers.
Versions
BenchCAD 1.0 is the released benchmark and the version every score on this page uses. BenchCAD 2.0 is in progress. It adds agentic evaluation with tool use, execution feedback and multi-turn refinement, and the BenchCAD site describes it as real-world engineering scenarios with about 100 public and 100 private cases. The private cases will give the benchmark a held-out set that is not on the public internet.
The data is CC-BY-4.0 on Hugging Face and the code is MIT on GitHub.
The official leaderboard has two settings. No-tools is a single response. Tools gives the model a Python sandbox where it can render, measure and iterate. Self-reported rows carry an asterisk meaning "voxel IoU only, not re-graded." Submissions made through a GitHub issue get re-graded, and the maintainers list their own runs separately.
Anthropic's harness
Anthropic runs Vision2Code on a random 1,000-file subset of the 17,900 files, which it says is "indicative of scores on the full set within a 0.01 voxel IoU margin." Each score is the mean of five runs with a 95% confidence interval, and every headline run uses adaptive thinking at max effort. With-tools runs get a container holding the image files and standard libraries, plus an image-cropping tool.
Three system cards document the harness and its fixes, and the fixes moved scores a lot:
- Fable 5 card, 9 June 2026. The headline no-tools run covered 17,874 files, dropping 26 records whose code produced no STEP file. Tools runs used a 1,000-file subset of those.
- Fable 5.1 card, 1 September 2026. Anthropic fixed a typo in the reference prompt that swapped all four camera positions. Grading now accepts raw shapes as well as Workplanes, after GPT-5.5 output errored on geometry that was equivalent. The parser now reads the last code fence instead of the first. The first two fixes are merged into the reference GitHub repository.
- Opus 5.5 card, 22 September 2026. Anthropic's implementation had been feeding the model the Hugging Face pre-rendered views at 128 x 128 px. The GitHub reference renders each view at 256 x 256 px, four times the pixels. Anthropic reproduced GPT-5.6 Sol's no-tools score at 256 px and could not at 128 px, then switched to 256 px and re-ran Fable 5.1 and Opus 5.
OpenAI's harness
OpenAI has not published its BenchCAD settings. Reasoning effort, scaffold, run count and tools are absent from the GPT-6 Astra system card and the GPT-6.1 Sol safety addendum, and neither document has a BenchCAD section. The Fable 5.1 card says GPT-5.6 Sol's scores were "publicly reported by OpenAI, evaluated on the full 17,900 files." Google's Gemini 4 methodology page has no BenchCAD entry.
Current leaderboard
Four groups of results exist, and they do not merge into one ranking. Anthropic's own runs share a harness and compare with each other. OpenAI's numbers come from a different file set with undisclosed settings. The BenchCAD maintainers have run only two frontier models themselves, both Grok. No Anthropic or OpenAI score has been independently re-graded.
Anthropic harness, with tools
Comparable within this table: same lab, same 1,000-file subset, 256 px views, five-run mean, max effort. Not comparable to any other table on this page.
| Model | Organization | Score (voxel IoU) | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Sonnet 5.5 | Anthropic | 0.963 | Container plus crop tool, adaptive thinking, max effort, 256 px | 2026-09-28 | Sonnet 5.5 system card, section 8.13.2. Vendor-reported. The 0.001 lead over Opus 5.5 is inside the 95% CI. |
| Claude Opus 5.5 | Anthropic | 0.962 | Container plus crop tool, adaptive thinking, max effort, 256 px | 2026-09-22 | Opus 5.5 system card, section 8.13.2. Vendor-reported. |
| Claude Fable 5.1 | Anthropic | 0.926 | Container plus crop tool, adaptive thinking, max effort, 256 px | 2026-09-22 | Opus 5.5 card, 8.13.2. Re-run at 256 px; the 128 px score was 0.843. Vendor-reported. |
| Claude Opus 5 | Anthropic | 0.899 | Container plus crop tool, adaptive thinking, max effort, 256 px | 2026-09-22 | Opus 5.5 card, 8.13.2. Re-run at 256 px; the 128 px score was 0.821. Vendor-reported. |
Anthropic harness, no tools
Comparable within this table and with the no-tools column of the historical table only where the resolution matches. Not comparable to the with-tools tables.
| Model | Organization | Score (voxel IoU) | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Sonnet 5.5 | Anthropic | 0.747 | Single response, adaptive thinking, max effort, 256 px | 2026-09-28 | Sonnet 5.5 card, 8.13.2 text and Fig 8.13.2.A. Vendor-reported. |
| Claude Opus 5.5 | Anthropic | 0.730 | Single response, adaptive thinking, max effort, 256 px | 2026-09-22 | Opus 5.5 card, 8.13.2. Vendor-reported. |
| Claude Fable 5.1 | Anthropic | 0.606 | Single response, adaptive thinking, max effort, 256 px | 2026-09-22 | Opus 5.5 card, 8.13.2. Re-run at 256 px; the 128 px score was 0.437. Vendor-reported. |
| Claude Opus 5 | Anthropic | 0.497 | Single response, adaptive thinking, max effort, 256 px | 2026-09-22 | Opus 5.5 card, 8.13.2. Re-run at 256 px; the 128 px score was 0.366. Vendor-reported. |
Without tools, Sonnet 5.5 leads Opus 5.5 by 0.017. With tools, the two tie. Tools add more than 0.2 IoU to every model in the table: 0.216 for Sonnet 5.5, 0.232 for Opus 5.5, 0.320 for Fable 5.1 and 0.402 for Opus 5. The weaker the model, the more it gains from being able to render and check its own output, and with tools the top three models sit within 0.037 of each other against a 0.141 spread without them.
The Sonnet 5.5 card's Figure 8.13.2.A also prints bar labels for Claude Sonnet 5 of 0.322 without tools and 0.519 with tools. No card text states the resolution or run details for those bars, so they stay out of the tables and the chart.
OpenAI-reported
Comparable within this table only. Not comparable to the Anthropic tables: OpenAI ran the full 17,900 files with undisclosed effort, scaffold and run count.
| Model | Organization | Score (voxel IoU) | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | 0.834 | With tools, full 17,900 files, settings not published | Cited in Fable 5.1 card, 2026-09-01 | Bar label in Sonnet 5.5 card Fig 8.13.2.A, "as publicly reported by OpenAI." Vendor-reported by OpenAI, not re-graded. |
| GPT-5.6 Sol | OpenAI | 0.706 | No tools, full 17,900 files, settings not published | Cited in Fable 5.1 card, 2026-09-01 | Same source. Anthropic reproduced this score in its own harness at 256 px, per the Opus 5.5 card. |
The same Anthropic figure prints a GPT-6 Astra with-tools bar of 0.959, also captioned as publicly reported by OpenAI. OpenAI's Astra announcement page returned HTTP 403 and was not read, and no Astra no-tools score, effort setting, cost or independent run exists, so Astra is not ranked here.
BenchCAD maintainers, independent runs
Comparable within this table only. Not comparable to the Anthropic or OpenAI tables: the maintainers use their own scorer and harness, and the page does not state the split or attempt budget.
| Model | Organization | Score (voxel IoU) | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Grok 4.6 | SpaceXAI | 0.8055 | Tools setting (Python sandbox), xhigh effort | Listed as of 2026-10-06 | benchcad.com/leaderboard. Run by the BenchCAD maintainers. |
| Grok 4.5 | SpaceXAI | 0.7771 | Tools setting (Python sandbox), high effort | Listed as of 2026-10-06 | benchcad.com/leaderboard. Run by the BenchCAD maintainers. |
The leaderboard labels the Anthropic and OpenAI figures "vendor self-reported voxel IoU, not re-graded."
Score against cost
The Sonnet 5.5 card plots an effort sweep for the with-tools runs: five runs per effort level, cost billed at API list prices with each run's actual cache usage. The card prints no data table. The costs below were read off a log axis and carry about 15% error either way. The three top scores are exact values from the card text; the rest were read from the figure to about 0.01. The figure does not label effort levels, so the labels name position on the curve.
With tools, score climbs from about 0.54 to 0.963 across a cost range of about 240x, from $0.025 to $6 per task. Sonnet 5.5 holds most of the frontier: its low, mid and top points are all on it, with Opus 5.5's mid point at about 0.76 for $0.5 between them. At the top, Sonnet 5.5 and Opus 5.5 land on the same score at the same cost. Fable 5.1 sits to the right and below. Its $10 point costs more than Sonnet 5.5's top point and scores 0.037 lower. Its $5 point draws on the frontier line as read, but $5 against $6 is inside the read error, and it scores 0.13 below Sonnet 5.5's top point.
Historical progression
All Anthropic rows are vendor-reported voxel IoU on Anthropic's harness. The resolution column decides what compares with what: 128 px rows compare with each other, 256 px rows compare with each other, and a 128 px score next to a 256 px score says nothing about model progress.
| Date | Model | No tools | With tools | Files and resolution | Source |
|---|---|---|---|---|---|
| 2026-06-09 | Claude Mythos 5 | 0.384 | 17,874 files; 128 px per the Opus 5.5 card | Fable 5 card, 8.16.4 | |
| 2026-06-09 | Claude Mythos Preview | 0.355 | 17,874 files; 128 px | Fable 5 card, 8.16.4 | |
| 2026-06-09 | Claude Opus 4.8 | 0.273 | 17,874 files; 128 px | Fable 5 card, 8.16.4 | |
| 2026-06-09 | Claude Mythos 5 | 0.379 | 0.650 | 1,000-file subset; 128 px | Fable 5 card, 8.16.4 |
| 2026-06-09 | Claude Mythos Preview | 0.356 | 0.610 | 1,000-file subset; 128 px | Fable 5 card, 8.16.4 |
| 2026-09-01 | Claude Fable 5.1 | 0.437 | 0.843 | 1,000-file subset; 128 px | Fable 5.1 card, 8.14.2 |
| 2026-09-01 | Claude Fable 5 | 0.376 | 0.675 | 1,000-file subset; 128 px | Fable 5.1 card, 8.14.2 |
| 2026-09-01 | Claude Opus 5 | 0.366 | 0.821 | 1,000-file subset; 128 px | Fable 5.1 card, 8.14.2 |
| 2026-09-22 | Claude Opus 5.5 | 0.730 | 0.962 | 1,000-file subset; 256 px | Opus 5.5 card, 8.13.2 |
| 2026-09-22 | Claude Fable 5.1 | 0.606 | 0.926 | 1,000-file subset; 256 px (re-run) | Opus 5.5 card, 8.13.2 |
| 2026-09-22 | Claude Opus 5 | 0.497 | 0.899 | 1,000-file subset; 256 px (re-run) | Opus 5.5 card, 8.13.2 |
| 2026-09-28 | Claude Sonnet 5.5 | 0.747 | 0.963 | 1,000-file subset; 256 px | Sonnet 5.5 card, 8.13.2 |
The Fable 5 card does not state the resolution of its runs. The Opus 5.5 card says all prior Claude runs used 128 px views, which is the basis for that column.
The re-runs show how much of the apparent jump came from the harness. Resolution alone moved Fable 5.1 by 0.169 without tools and 0.083 with tools, and Opus 5 by 0.131 and 0.078.
Paper-era results
The paper ran its own Vision2Code protocol in May 2026 and reports a combined total score next to IoU. These numbers do not compare with any Anthropic, OpenAI or maintainer row above.
| Model | Voxel IoU | Total score | Notes |
|---|---|---|---|
| Qwen3-VL-2B, RL on BenchCAD (iid) | 0.752 | 0.768 | Fine-tuned on in-distribution families |
| CADEvolve v3 | 0.750 | 0.601 | CAD specialist, about 2B parameters |
| Gemini 3.1 Pro (thinking) | 0.279 | 0.397 | General frontier model |
| Claude Opus 4.7 (no thinking) | 0.274 | 0.383 | General frontier model |
| Claude Opus 4.7 (thinking) | 0.267 | 0.378 | General frontier model |
| GPT-5.3 (thinking) | 0.207 | 0.212 | General frontier model |
Source: BenchCAD paper, Table 10.
In May, a 2B model fine-tuned on BenchCAD beat every general frontier model by 0.47 IoU or more. The other paper tasks tell a different story. On Code Edit (Table 3), GPT-5.3 thinking scored 0.865, Claude Opus 4.7 thinking 0.853 (0.811 without thinking), Gemini 3.1 Pro thinking 0.837, GPT-5.3 without thinking 0.740, o3 0.708 and GPT-4o 0.615. Frontier models could already edit CAD code well. They could not build it from pictures.
Vision QA (Table 2) shows where the gap lives. Gemini 3.1 Pro scored 0.587 total (0.576 with thinking), Claude Opus 4.7 0.526 (0.530 with thinking) and GPT-5.3 chat 0.513 (0.514 with thinking), against a blank-image baseline of 0.375. The best Code QA models reach about 0.838 on the same questions. Opus 4.7 scored 0.699 on L1 holistic recognition, so the models could tell what a part was and lost their points on the numeric detail.
Documented failure modes
The paper catalogs specific errors in section 5.2 and Figure 6:
- Fine structure goes missing. Models get the outer shape and drop threads and chamfers.
- Wrong base plane. A part extruded from XZ comes back extruded from XY.
- Twist-extrusion disappears. A twisted bracket becomes two perpendicular brackets.
- Standards get flattened. The DIN 2095 spring's closed-and-ground ends turn into a uniform helix.
The hard tier is a cliff. Pass rate collapses on the 26 hard-tier families that use helical sweeps, twist-extrusion and lofted Booleans, with a gap of more than 45 points from the easy tier for the strongest reasoning model (Table 12). These are the same operations the abstract says models replace with sketch-and-extrude.
Vision is the bottleneck, not code. Vision QA tops out at 0.587 while Code QA reaches about 0.838 on matched questions, which the paper calls a "Holistic Spatial and Detailing Deficit." The tool results line up with that. Anthropic says Claude's performance scales with test-time compute, "particularly when the models are equipped with tools that enable visual verification of intermediate outputs." In Anthropic's tables, tools add 0.216 to 0.402 IoU, the largest single lever on this benchmark.
Fine-tuning does not fully generalize. The paper's Qwen3-2B RL model drops from 0.752 IoU on in-distribution families to 0.714 on held-out families (Table 10).
The harness moves scores more than most model releases do. A fourfold increase in render resolution moved Fable 5.1 from 0.437 to 0.606 without tools. Before that, Anthropic's implementation had a camera-position typo, rejected valid raw-shape outputs and parsed the wrong code fence. Any BenchCAD number without its render size and grading code version attached is hard to use.
The benchmark's scope is narrower than its name. Standard-anchored parts are not checked for standard compliance, since the paper does not validate tolerances. The paper says the parts "do not capture the full diversity of proprietary industrial design," and they carry no manufacturing tolerances, undocumented design intent or assembly information. The inputs are clean synthetic renders, not photos or scans.
Contamination is unchecked. All 17,900 programs, renders and code have been public under CC-BY-4.0 since May 2026, and no lab has published a decontamination check for BenchCAD. BenchCAD 2.0's private cases are the first fix on the way.
How BenchCAD compares to related benchmarks
MMMU asks multiple-choice questions about images, so a model can guess its way to points. BenchCAD requires a program that runs and produces geometry that gets voxel-compared, so guessing cannot score.
SWE-bench also grades code by execution, but it scores repository patches pass or fail against unit tests. BenchCAD scores a fresh program on a continuous scale, so partial credit is the norm. A BenchCAD score of 0.9 does not mean 90% of parts are correct. It means the average part overlaps its reference by 90%.
Chartography is the closest sibling. Anthropic reports it in the same card section (8.13.1 of the Sonnet 5.5 card), with and without tools and with an effort-versus-cost sweep. Both turn an image into structured output. Chartography scores accuracy and BenchCAD scores voxel IoU, so the two numbers do not share a scale.
CAD training corpora such as CADEvolve are a different kind of thing. The paper's CADEvolve v3 specialist, at about 2B parameters, scored 0.750 IoU on Vision2Code against 0.2 to 0.28 for frontier models under the same protocol. BenchCAD is an evaluation set built to span 106 families, 47 standard codes and 49 operations, not a training corpus.
For the broader picture of how vendors run and report these evals, see LLM benchmarks and LLM evaluation. Sibling pages cover OSWorld 2 and Terminal-Bench 4.
What BenchCAD means for teams choosing a model
The numbers here are Anthropic's harness with tools unless stated, costs are read from the Sonnet 5.5 card figure, and list prices come from Anthropic's pricing page.
Pick Sonnet 5.5 or Opus 5.5 at the top, and do not expect Sonnet to be cheaper there. They tie at 0.963 and 0.962, and both cost about $6 per task. Sonnet 5.5 lists at $2 input and $10 output per million tokens against $4 and $20 for Opus 5.5, but at top effort the lower list price does not show up in cost per task. Below the top, Sonnet 5.5's mid point is the best score under $0.25 (about 0.70 at $0.2), and Opus 5.5's mid point reaches about 0.76 for $0.5.
Skip Fable 5.1 for this work. It lists at $10 and $50 per million tokens and scores 0.926 at about $10 per task, against 0.963 at about $6 for Sonnet 5.5.
Give the model tools. Tools raise every model by 0.2 IoU or more. Without tools Sonnet 5.5 tops out at 0.747. Even the cheapest with-tools point, about 0.54 for $0.025, costs roughly 1/40 of the no-tools top-effort run (about $1, per the figure's dashed lines), though it scores below it. If your CAD pipeline can render a candidate and hand the image back to the model, build that loop before you spend on a bigger model.
Treat OpenAI's numbers as unverified. GPT-5.6 Sol reports 0.706 without tools and 0.834 with tools on the full file set with undisclosed settings. That with-tools score sits below every model in the Anthropic with-tools table, but the harnesses differ and nobody has re-graded it. OpenAI has published no cost per task.
Discount older Claude BenchCAD numbers. Anything published before the Opus 5.5 card on 22 September 2026 is low by 0.078 to 0.083 with tools and 0.131 to 0.169 without, because Anthropic rendered views at 128 px instead of the reference 256 px. When you run your own evals, pin the render size at 256 x 256 and record the file subset.
BenchCAD's inputs are clean synthetic renders, so if your work starts from real drawings or scans, run the same kind of eval on your own parts in Klu. Compare models on the Klu LLM leaderboard, and see the engineering page for evals on engineering workflows.