Agents' Last Exam (ALE)

A UC Berkeley RDI benchmark that tests whether an AI agent can finish long professional tasks in real software, with deterministic checkers grading the deliverable file

Computer usePass rate, score

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 15 min read

What is Agents' Last Exam?

Agents' Last Exam (ALE) tests whether an AI agent can finish a long piece of paid professional work in real software and hand back a file that a program can check. The agent gets a task description, some input files and a sandbox with the target application already installed. It might need to set up a toolpath in Siemens NX, build a scene in Unreal Engine, run a mold-flow study in Moldex3D or clean up a brain scan in FSLeyes. Then it writes its deliverable to an output folder, and a grader compares that file against a hidden reference.

UC Berkeley RDI built it. The leaderboard page says the project is "co-led by UC Berkeley RDI and the RDI Foundation," and Snorkel AI hosts a mirror of the board. The paper, led by Yiyou Sun with 310 listed authors, went up on June 3, 2026, with a v2 on June 11. It reports 1,490 task instances drawn from 960 unique workflows, sorted into 13 industry clusters and 55 subfields that map to the US O*NET / SOC 2018 occupational taxonomy. The tasks come from practitioners' completed projects, and between 250+ and 300+ experts took part. The paper abstract says 250+, and the leaderboard page and the paper body say 300+.

The paper places ALE against three groups of earlier benchmarks. Knowledge tests measure what a model knows, not what it can do. Agentic suites such as OSWorld and WebArena cover narrow domains with tasks curators wrote themselves. Economic evaluations such as GDPval and RLI use real work but cover 16 and 14 of the 55 industries, and they need human graders. ALE covers all 55 subfields and grades with code. In the open-sourced workflows, 93.2% use deterministic judges and 6.8% use an LLM judge.

The scope is software-mediated work. The abstract calls it "non-physical industries," and the domains run through manufacturing, biomolecular design, animation, robotics, agriculture, finance, law and education. The leaderboard lists After Effects, Siemens NX, Unreal Engine, Moldex3D, Rhino 3D and FSLeyes "and 49 more applications."

ALE is not a GUI-grounding test and not a test of how agents cope with OS popups or a drifting desktop. That is OSWorld 2 and ScreenSpot-Pro territory. ALE grades the work product.

How the tasks and harness work

Every task follows the same directory contract, per the repo README and the paper. input/ holds read-only files, software/ holds the pre-installed application, output/ is the only folder the agent can write to, and reference/ holds the hidden ground truth. The agent sees a natural-language description and a deliverable spec. The rubric and the reference stay out of reach.

Tasks run in Linux or Windows VMs and containers. The supported sandbox providers include Google Cloud, AWS, local Docker and QEMU/KVM. Agent frameworks run as shipped: Codex, Claude Code, Cursor CLI, OpenClaw, ALE-Claw, Droid, Gemini CLI, Grok CLI, Hermes, ForgeCode, Terminus, OpenHands and others. Two open harnesses ship with the framework, the official Claude Code CLI and the in-tree OpenClaw harness. ALE-Claw is the reference harness.

The paper also adds a mode it calls GUI-as-Tool. It exposes 14 desktop actions over an MCP bridge next to ordinary shell commands, so one model can reason over shell output and screenshots in the same loop. This matters because 34% of public instances name graphical software as the primary tool. A CLI-only agent cannot do those tasks the intended way.

Each run has a five-hour wall-clock cap. When the cap hits, the harness stops the agent and the grader scores whatever is in output/. The paper states no step limit. In the paper's main runs, 3.8% of runs hit the cap, about 1% for lightweight harnesses and 5.7% for OpenClaw. Capped runs averaged a score of 20.7 against 33.2 for runs that finished earlier.

The paper's cost estimate for a single run is "$3-10 on average," taking "tens of minutes to hours" per task. That puts one ALE task in a different weight class from a coding benchmark item.

How scoring works

Graders are deterministic wherever the output allows it: exact or hash matches, tolerance checks on numeric and tabular fields, geometric distance, world-state checks, and narrow yes/no probes from a vision model. Free-text rubrics appear only when nothing deterministic fits. The leaderboard sums it up as "Hidden references plus deterministic graders, not LLM-as-a-judge."

Scores combine a gate and a score. Each task has binary gates, such as "the file parses" or "the toolpath is collision-free." Fail a gate and the task scores zero, however good the rest of the work looks. Pass the gates and the graded checks add up to partial credit.

The board reports two numbers per row:

  • Pass Rate is the share of tasks fully completed. It is the headline metric.
  • Score is the average graded outcome, partial credit included.

The two numbers sit far apart. The Overall leader posts a 38.2% pass rate and a 63.2 score. An agent that gets most of a deliverable right still fails the task if one required piece is missing, and Pass Rate is the number to read if your workflow has no use for a nearly finished file.

The paper reports three-run standard deviations of roughly ±1 to ±3 on score for a subset of configurations. Most configurations have one run.

Versions, tiers and tabs

The live board is labeled ALE-V1, snapshot dated September 29, 2026. The paper's main runs come from the GPT-5.5, Claude Opus 4.7 and Claude Fable 5 generation. The board adds GPT-5.6, GPT-6, Claude Opus 5 and 5.5, Kimi K3, Grok 4.5 and Muse Spark 1.3.

About 150 tasks are public. The sources give three counts: the Snorkel page says 147, the repo README says "around 150," and the paper says both 150 public instances and "the full 152-task public set." In the paper's split, 1,017 instances are private and 323 were pending quality control. Held-out private tasks score the official leaderboard, and the public pool rotates. Snorkel's page says: "Every ~6 months, a new public subset releases with fresh instances. Private tasks rotate into the public pool, retired public tasks rotate out."

The public set is harder than the full pool. The paper ran Claude Code with Opus 4.7 on everything and got a higher pass rate on the full pool, because the public set holds the entire hardest tier while the private pool leans toward easier tasks. Per-cluster pass rates still track closely between the two, at a Pearson r of 0.89.

The paper defines three difficulty tiers plus a CLI subset:

TierPaper task countBoard tab description
Near-Term67"partially solvable today"
Full-Spectrum55, one per subfield"50+ industries"
Last-Exam38, the hardest"frontier difficulty"
ALE-CLI105 Linux-only"linux task only"

The board has five tabs: Overall plus those four. Each tab has its own rows and picks its own best effort setting per model. The board does not say how many tasks sit behind each tab or whether they match the paper's counts. Most Near-Term pass rates fit a count out of 67, but two Codex GPT-5.5 values do not, so this page does no per-task division. Costs and runtimes below are the board's published totals for each tab.

Current leaderboard

Every row below comes from the Snorkel mirror, whose embedded data names the canonical board as its source, with an as-of date of 2026-09-29 and version ALE-V1. The canonical page loads its rows by script, so none were read from it directly. The board publishes no per-row evaluation date and no field saying who ran each row, whether a vendor submitted it, or whether anyone verified it. That holds for every row here, including rows where the harness vendor and the model vendor are the same company. No lab system card or release post retrieved for this page reports an ALE score.

Three more things to know before reading the numbers:

  • Costs drift. The canonical site recalculates cost estimates without rerunning evaluations. GPT-5.6 Sol's Overall cost moved from $762 to $772, and GPT-6 Astra's cost and token counts were republished on September 29 with pass rate, score and runtime unchanged. The board does not publish how it estimates cost.
  • Claude Fable 5 rows carry a flag. The source site states that "the variant served during evaluation may differ from the model's full-capability tier, and re-runs cannot guarantee the higher-tier variant is selected. Scores may understate the model's true ceiling."
  • Runtime changed method. The source site updated its runtime methodology on July 10, 2026, so runtimes do not line up with earlier snapshots.

Cost and runtime are totals for the tab's task set. Effort is the board's best-effort pick, meaning the setting with the highest pass rate, ties broken by score.

On the Overall tab, Claude Code with Claude Opus 5.5 at Max effort leads at 38.2% pass and a 63.2 score. Codex with GPT-6 Astra at Max is second at 34.2% and 59.3, 4.0 points behind on pass rate. Those are model-plus-harness pairs. The board has no Opus 5.5 run in Codex and no Astra run in Claude Code, so the gap between the two models alone is unmeasured.

Overall tab, Codex harness

Same harness and tab for every row, best-effort rows only. Comparable within this table. Comparing to the Claude Code table compares model-plus-harness pairs, not models.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI34.2% pass, 59.3Codex, Max2026-09-29 snapshotSnorkel mirror. $1,100, 146h16m, 580.6M in / 6.3M out tokens
GPT-6 SolOpenAI32.2% pass, 56.5Codex, XHigh2026-09-29 snapshotSnorkel mirror. $318, 123h53m, 929.8M / 5.8M tokens
Muse Spark 1.3Meta32.2% pass, 55.8Codex, XHigh2026-09-29 snapshotSnorkel mirror. $428, 80h37m, 2.1B / 11.8M tokens
GPT-5.6 SolOpenAI30.6% pass, 53.6Codex, XHigh2026-09-29 snapshotSnorkel mirror. $772, 94h39m, 762.9M / 3.8M tokens
GPT-5.6 LunaOpenAI30.3% pass, 49.4Codex, XHigh2026-09-29 snapshotSnorkel mirror. $235, 66h07m, 1.4B / 4.2M tokens
GPT-5.6 TerraOpenAI28.0% pass, 50.7Codex, Max2026-09-29 snapshotSnorkel mirror. $545, 118h39m, 1.2B / 5.8M tokens
GPT-5.5OpenAI26.6% pass, 47.9Codex, XHigh2026-09-29 snapshotSnorkel mirror. $602, 97h08m, 560.5M / 6.4M tokens
GPT-6 LunaOpenAI25.0% pass, 50.9Codex, Max2026-09-29 snapshotSnorkel mirror. 131h50m, 1.2B / 12.1M tokens. Cost cell omitted here and from every cost comparison
ALE leaderboard via Snorkel mirrorCodexOverall tab, ALE-V1Pass rate, scoresnorkel.ai

Astra leads the Codex field by 2.0 points over GPT-6 Sol and Muse Spark 1.3. Sol gets there for $318 against Astra's $1,100. Meta's Muse Spark 1.3 is the only non-OpenAI model in this table, and it ties Sol on pass rate with the shortest runtime of the top four.

Overall tab, Claude Code harness

Same harness and tab for every row, best-effort rows only. Comparable within this table. Not comparable model-to-model with the Codex table.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 5.5Anthropic38.2% pass, 63.2Claude Code, Max2026-09-29 snapshotSnorkel mirror. $1,340, 182h45m, 4.2B / 25.9M tokens. Listed at Max only
Claude Opus 5Anthropic32.2% pass, 55.9Claude Code, High2026-09-29 snapshotSnorkel mirror. $1,108, 133h30m, 1.1B / 11.9M tokens
Qwen3 8-MaxAlibaba27.0% pass, 52.5Claude Code, XHigh2026-09-29 snapshotSnorkel mirror. $486, 202h41m, 1.2B / 12.4M tokens
Kimi K3Moonshot27.0% pass, 50.7Claude Code, Max2026-09-29 snapshotSnorkel mirror. $464, 215h32m, 950.6M / 7.4M tokens
Claude Opus 4.8Anthropic27.0% pass, 45.1Claude Code, Max2026-09-29 snapshotSnorkel mirror. $3,985, 179h18m, 1.7B / 22.1M tokens
Claude Fable 5Anthropic25.7% pass, 48.7Claude Code, XHigh2026-09-29 snapshotSnorkel mirror. $4,340, 70h00m, 2.7B / 10.4M tokens. Source-site variant flag applies. Adaptive row: 22.0% / 40.5 / $2,315 / 106h14m
GLM-5.2Zhipu20.4% pass, 41.1Claude Code, Max2026-09-29 snapshotSnorkel mirror. $1,086, 107h42m, 1.3B / 7.7M tokens
Seed 2.1 ProByteDance19.5% pass, 41.9Claude Code, effort not shown2026-09-29 snapshotSnorkel mirror. $936, 155h14m, 2.1B / 8.4M tokens
Claude Opus 4.7Anthropic13.8% pass, 35.8Claude Code, High2026-09-29 snapshotSnorkel mirror. $1,793, 42h36m, 456.4M / 3.7M tokens
ALE leaderboard via Snorkel mirrorClaude CodeOverall tab, ALE-V1Pass rate, scoresnorkel.ai

Opus 5.5 is 6.0 points clear of Opus 5 here. Below them, Qwen3 8-Max, Kimi K3 and Opus 4.8 tie at 27.0%, and the first two get there for about an eighth of Opus 4.8's $3,985.

Near-Term tab: pass rate against estimated tab costTop eight best-effort rows on the Near-Term tab, Codex unless noted
45%50%55%$50$100$200$500
FrontierOpenAIAnthropicBehind the frontier
Est. cost for the Near-Term tab (USD), log scale, lower to the right
Source: Snorkel mirror of the canonical ALE board, Near-Term tab, snapshot 2026-09-29, read 2026-10-06. Cost is the board's estimated total for the whole tab, not per task; the board does not publish its cost method or the tab's task count. Codex GPT-5.5 (44.0%, a value that does not fit a count over 67) and GPT-6 Luna (a $6 cost outlier) are left off.

The chart uses the Near-Term tab because it is the tier where today's agents finish about half the work, so the spread between rows means something. Four rows sit on the cost frontier: GPT-5.6 Luna at $69 and 49.3%, GPT-6 Sol at $81 and 50.7%, GPT-6 Astra at $283 and 52.2%, and Opus 5.5 at $480 and 53.7%. GPT-6 Sol gets within 1.5 points of Astra for 29% of the cost. Opus 5.5 adds another 1.5 points over Astra for 1.7x the money. Opus 5, GPT-5.6 Sol, Muse Spark 1.3 and GPT-5.6 Terra all sit behind the frontier.

Effort ladders in the Overall tab

Same model, harness and tab within each block; compare across effort settings only. Source is the Snorkel mirror's per-row effort data alone.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-6 AstraOpenAI34.2% pass, 59.3Codex, Max2026-09-29 snapshotSnorkel mirror. $1,100, 146h16m
GPT-6 AstraOpenAI32.2% pass, 58.3Codex, XHigh2026-09-29 snapshotSnorkel mirror. $898, 96h23m
GPT-6 AstraOpenAI31.6% pass, 57.8Codex, High2026-09-29 snapshotSnorkel mirror. $783, 90h30m
GPT-6 AstraOpenAI32.2% pass, 57.6Codex, Medium2026-09-29 snapshotSnorkel mirror. $733, 72h42m
GPT-6 AstraOpenAI29.6% pass, 53.4Codex, Low2026-09-29 snapshotSnorkel mirror. $540, 65h38m
GPT-6 SolOpenAI31.6% pass, 58.6Codex, Max2026-09-29 snapshotSnorkel mirror. $428, 160h46m
GPT-6 SolOpenAI32.2% pass, 56.5Codex, XHigh2026-09-29 snapshotSnorkel mirror. $318, 123h53m
GPT-6 SolOpenAI31.6% pass, 55.3Codex, High2026-09-29 snapshotSnorkel mirror. $321, 122h34m
GPT-6 SolOpenAI28.9% pass, 53.8Codex, Medium2026-09-29 snapshotSnorkel mirror. $292, 125h38m
GPT-6 SolOpenAI27.6% pass, 52.4Codex, Low2026-09-29 snapshotSnorkel mirror. $153, 89h27m
Claude Opus 5Anthropic32.2% pass, 55.9Claude Code, High2026-09-29 snapshotSnorkel mirror. $1,108
Claude Opus 5Anthropic30.9% pass, 52.7Claude Code, Max2026-09-29 snapshotSnorkel mirror. $1,483
Claude Opus 5Anthropic30.3% pass, 55.5Claude Code, XHigh2026-09-29 snapshotSnorkel mirror. $1,523
Claude Opus 5Anthropic28.9% pass, 53.0Claude Code, Medium2026-09-29 snapshotSnorkel mirror. $804
Claude Opus 5Anthropic27.6% pass, 51.9Claude Code, Low2026-09-29 snapshotSnorkel mirror. $617
ALE leaderboard via Snorkel mirrorOverall tab, ALE-V1Pass rate, scoresnorkel.ai

More test-time compute helps Astra and stops helping the other two. Astra's Max beats its Low by 4.6 points for 2.0x the cost. GPT-6 Sol peaks at XHigh, and its Max costs $110 more for 0.6 fewer points. Opus 5 peaks at High. Its Max and XHigh both cost more and pass fewer tasks.

Other harnesses on the Overall tab

One row per harness, each with a different model; this is a reference list, not a ranking. Rows do not rank against each other or against the Codex and Claude Code tables.

ModelOrganizationScoreHarness and setupDateSource and notes
Kimi K3Moonshot28.3% pass, 51.6Kimi Code, Max2026-09-29 snapshotSnorkel mirror. $606, 186h05m. 27.0% in Claude Code
Grok 4.5xAI27.0% pass, 51.2Grok Build, High2026-09-29 snapshotSnorkel mirror. $337, 89h08m
Agnes-2.5-Pro-BetaNot published21.7% pass, 42.7Agnes Harness, effort not reported2026-09-29 snapshotSnorkel mirror. No cost published, 219h12m
Gemini 3.1 ProGoogle16.4% pass, 32.7Gemini CLI, High2026-09-29 snapshotSnorkel mirror. $2,018, 92h21m
Grok 4.3xAI7.2% pass, 21.3Grok CLI, effort not shown2026-09-29 snapshotSnorkel mirror. $285, 43h07m
ALE leaderboard via Snorkel mirrorOverall tab, ALE-V1Pass rate, scoresnorkel.ai

The board has no Gemini 4 entry, and no source retrieved for this page reports a Gemini 4 ALE result.

Same model, different harness

Rows share a model and the High effort setting; only the harness differs. Comparable within each model's block. The Codex GPT-5.5 High row is not listed because the board's best-effort row for that pair is XHigh.

ModelOrganizationScoreHarness and setupDateSource and notes
GPT-5.5OpenAI23.0% pass, 45.8ALE-Claw, High2026-09-29 snapshotSnorkel mirror. $310, 35h34m
GPT-5.5OpenAI21.7% pass, 42.1OpenClaw, High2026-09-29 snapshotSnorkel mirror. $447, 82h38m
GPT-5.5OpenAI19.7% pass, 39.2Droid, High2026-09-29 snapshotSnorkel mirror. $242, 63h49m
Claude Opus 4.7Anthropic21.1% pass, 42.5Cursor CLI, High2026-09-29 snapshotSnorkel mirror. $1,899, 63h49m
Claude Opus 4.7Anthropic18.4% pass, 40.5ALE-Claw, High2026-09-29 snapshotSnorkel mirror. $1,132, 77h39m
Claude Opus 4.7Anthropic15.8% pass, 35.2OpenClaw, High2026-09-29 snapshotSnorkel mirror. $1,689, 124h10m
Claude Opus 4.7Anthropic13.8% pass, 35.8Claude Code, High2026-09-29 snapshotSnorkel mirror. $1,793, 42h36m
Claude Opus 4.7Anthropic13.5% pass, 31.7Droid, High2026-09-29 snapshotSnorkel mirror. $1,356, 28h16m
ALE leaderboard via Snorkel mirrorOverall tab, ALE-V1, High effortPass rate, scoresnorkel.ai

Opus 4.7 spans 7.6 points across five harnesses, from 13.5% in Droid to 21.1% in Cursor CLI. Its own vendor's harness, Claude Code, comes second to last. GPT-5.5 spans 3.3 points, from 19.7% in Droid to 23.0% in ALE-Claw. The agent architecture wrapped around the model is part of the result.

Difficulty tiers

Each tab is a different task set, and each cell is that tab's best-effort row, so effort can change across a row. Compare within a column only. Cells read pass / score / estimated tab cost.

ModelHarnessNear-TermFull-SpectrumALE-CLILast-Exam
Claude Opus 5.5Claude Code53.7 / 83.9 / $480 (Max)32.7 / 58.4 / $533 (Max)34.3 / 63.7 / $855 (Max)15.8 / 33.9 / $380 (Max)
GPT-6 AstraCodex52.2 / 82.3 / $283 (XHigh)30.9 / 51.8 / $329 (Max)33.3 / 61.4 / $558 (Max)10.5 / 27.4 / $350 (High)
GPT-6 SolCodex50.7 / 80.9 / $81 (High)29.1 / 53.1 / $109 (Max)32.4 / 58.8 / $179 (High)13.2 / 26.7 / $186 (Medium)
Claude Opus 5Claude Code49.3 / 80.1 / $296 (High)29.1 / 47.9 / $346 (High)29.5 / 57.3 / $557 (High)13.2 / 19.7 / $654 (Max)
GPT-5.6 LunaCodex49.3 / 75.6 / $69 (XHigh)29.1 / 42.5 / $107 (Max)29.5 / 50.7 / $144 (XHigh)2.6 / 15.0 / $197 (Max)
Muse Spark 1.3Codex47.8 / 77.8 / $103 (XHigh)29.1 / 47.4 / $130 (Max)33.3 / 57.5 / $256 (XHigh)10.5 / 23.8 / $225 (XHigh)
GPT-5.6 SolCodex47.8 / 78.8 / $322 (Max)30.0 / 46.6 / $256 (XHigh)28.6 / 53.6 / $446 (XHigh)5.3 / 19.4 / $325 (XHigh)
Claude Opus 4.8Claude Code43.3 / 64.0 / $1,106 (Max)Not in top 1425.7 / 44.1 / $1,953 (Max)Not in top 14
GPT-6 LunaCodex38.8 / 73.4 / cost omitted (Max)Not in top 14Not in top 1410.5 / 23.2 / cost omitted (XHigh)
Claude Fable 5Claude Code37.3 / 71.1 / $1,018 (XHigh)25.5 / 42.0 / $1,826 (XHigh)23.8 / 48.5 / $1,663 (XHigh)7.9 / 16.5 / $2,091 (XHigh)
ALE leaderboard via Snorkel mirrorNear-Term, Full-Spectrum, ALE-CLI, and Last-Exam tabs, ALE-V1Pass / score / estimated tab cost, best effort per tabsnorkel.ai

Claude Code with Opus 5.5 is the top row on every tab. Codex with Astra is second on Overall, Near-Term and Full-Spectrum and ties Muse Spark 1.3 at 33.3% on ALE-CLI.

The Last-Exam column needs a different reading. Its pass rates are task counts out of 38, so 15.8% is six tasks, 13.2% is five, 10.5% is four and 2.6% is one. A single task moves the number 2.6 points. Opus 5.5's lead there is one task over GPT-6 Sol and Opus 5, and two over Astra. Other Last-Exam rows: Qwen3 8-Max in Claude Code at XHigh 10.5 / 24.1 / $217, Kimi Code with Kimi K3 10.5 / 20.6 / $308, Grok Build with Grok 4.5 10.5 / 23.7 / $115, and Claude Code with Kimi K3 7.9 / 20.5 / $200.

Historical results

ALE is four months old, so its history is one model generation inside a few harnesses. The paper's June runs and the September board are two vintages with different setups, and they stay in separate tables.

Paper runs, June 2026

From arXiv 2606.05405 Table 1, v2 dated June 11, 2026. Every harness had the GUI-as-Tool mode added. Full-pass rates; these are the paper authors' runs on older models and do not compare directly with the live board. The paper's per-row effort settings were not extracted for this page.

Harness and modelOverallNear-TermFull-SpectrumLast-Exam
Codex, GPT-5.524.0%38.1%22.7%0.0% (score 11.2)
ALE-Claw, GPT-5.523.0%32.8%23.6%2.6%
Claude Code, Fable 522.0%34.3%20.9%0.0% (score 5.2)
OpenClaw, GPT-5.521.1%35.8%18.2%0.0%
Cursor, GPT-5.520.7%32.1%20.0%2.6%
Cursor, Composer 2.520.4%34.3%18.2%0.0%
Cursor, Opus 4.720.4%29.9%20.0%2.6%
Droid, GPT-5.519.1%29.9%16.4%2.6%
ALE-Claw, Opus 4.718.4%28.4%18.2%0.0%
Gemini CLI, Gemini 3.1 Pro15.8%26.9%12.7%0.0%
Claude Code, Opus 4.815.8%26.9%10.9%0.0%
Claude Code, Opus 4.713.2%20.9%12.7%0.0%
Droid, Opus 4.712.8%27.6%3.6%0.0%
Codex, GPT-5.47.2%13.4%3.6%0.0%
Grok CLI, Grok 4.36.6%9.0%7.3%0.0%
ALE paper authorsEach harness with GUI-as-Tool mode addedFull-pass ratearxiv.org

The paper also held the harness fixed and swapped the model. In OpenClaw, 14 models ran from GPT-5.5 at 21.1% down to Grok 4.3 at 4.3%, with Opus 4.7 at 15.1, Gemini 3.1 Pro and Opus 4.6 at 14.1, DeepSeek V4 Pro at 12.4 and Sonnet 4.6 at 9.9. Appendix D.4 puts the effect of the model at roughly 3x the spread from harness choice among well-built systems. On the paper's ALE-CLI subset, Codex with GPT-5.5 scored 23.3%, and Claude Code and ForgeCode with Sonnet 4.6 scored 16.7% and 13.3%. The live ALE-CLI tab is a different run set.

Board series, September 2026

Each series fixes one harness and one effort setting. Different vintage and setup from the paper table; the board makes no GUI-as-Tool statement.

SeriesResults
Claude Code at Max effortOpus 4.8 27.0% (score 45.1), Opus 5 30.9% (52.7), Opus 5.5 38.2% (63.2)
Claude Code at High effortOpus 4.7 13.8% (35.8), Opus 5 32.2% (55.9)
Codex at XHigh effortGPT-5.5 26.6% (47.9), GPT-5.6 Sol 30.6% (53.6), GPT-6 Sol 32.2% (56.5), GPT-6 Astra 32.2% (58.3)
Codex at Max effortGPT-6 Sol 31.6% (58.6), GPT-6 Astra 34.2% (59.3)
Last-Exam tab, Claude Code at MaxOpus 5 13.2%, Opus 5.5 15.8% (6 of 38)

Two facts stand out. Claude Code with Opus 4.7 scored 13.2% in the paper and 13.8% on the board at High, so the two run sets agree within 0.6 points on the one configuration they share. And within Claude Code at High effort, Opus 4.7 to Opus 5 is an 18.4-point jump. The hardest tier moved too. The paper abstract says the "average full pass rate is below 1%" on Last-Exam across mainstream configurations, and its best was one task of 38. The board's best is now six.

Where agents fail

The paper analyzed failed Claude Code runs with Opus 4.7. About three quarters of failures were understanding and approach problems: the agent misread what the deliverable needed or picked the wrong way to make it. About a quarter were execution and skill problems. Appendix D.4 traces the main bottleneck to reasoning and domain knowledge. One habit shows up again and again. Agents write ad-hoc scripts instead of opening the domain software the task was built around.

GUI use is the clearest gap. A third of public tasks name a graphical application as the primary tool, and in most configurations the share of GUI actions stays small. Snorkel's reading-group post puts GUI evaluation at about 50x slower than CLI.

Results also vary by field. Averaged over harnesses for Fable 5 and GPT-5.5, computational mathematics and agriculture and environment score highest, about 55 to 85%. Business and legal land around 50 to 55%, and education is lowest, below 25%.

Spending more does not buy a better result. In Claude Code, Fable 5 spends $4,340 for a 48.7 score while Opus 5.5 spends $1,340 for 63.2. In the paper's cost analysis, ALE-Claw with GPT-5.5 reached a 45.8 mean score for $326, while ALE-Claw with Opus 4.7 cost 3.6x more for 40.5. Claude Code with Fable 5 spent $2,402, the most of any configuration, for the same 40.5. Cursor with Composer 2.5 was the cheapest, at 38.5 for $87. ALE-Claw with GPT-5.5 also finished in about 51 hours of wall-clock time against 466 hours for Claude Code with Opus 4.8.

Timeouts cost real points. Capped runs average 20.7 against 33.2 for others, and the Last-Exam tier times out at 6.4%, double the 3.1% and 3.2% of the easier tiers.

Contamination defense rests on the private pool and the six-month rotation. The roughly 150 public tasks are readable by anyone, training pipelines included. Neither the paper nor the leaderboard documents a contamination measurement.

Several things are not published: the cost estimation method, per-row evaluation dates, who ran each board row, the number of tasks behind the Overall tab, a step limit, and any vendor-reported ALE score. OpenAI's GPT-6 Astra model page and Anthropic's Opus 5.5 announcement list pricing and dates but no ALE result.

GDPval is the closest relative. Both grade professional deliverables, but GDPval relies on human expert graders, and the ALE paper counts its coverage at 16 of 55 industries against ALE's 55. ALE swaps the human grader for a deterministic checker, which makes it cheaper to rerun and harder to argue with. The cost is that it only works where a checker can be written.

Terminal and coding benchmarks look easy next to it. The ALE paper says Codex with GPT-5.5 "achieves 82% on Terminal-Bench," yet the same pair reached 23.3% on ALE-CLI and stayed under 50% on ALE's easiest tier. The paper does not name a Terminal-Bench version, so do not read that 82% against Terminal-Bench 4.0. SWE-bench is the test-graded coding analogue. It grades by execution too, but for one profession instead of 55.

OSWorld 2 tests GUI-only computer use on general desktop apps. ALE gives the agent a shell and a GUI and grades the professional file it produces, not the clicks. GAIA asks for short answers to general-assistant questions, while ALE wants multi-hour deliverables. For narrower workflow suites, see AutomationBench for business automation and the Harvey legal agent benchmark for law.

What it means for teams choosing a model

Pick the harness and the model together. On the same model at High effort, the harness moved Opus 4.7 by 7.6 points and GPT-5.5 by 3.3. The top two Overall rows, Claude Code with Opus 5.5 and Codex with Astra, have never run in each other's harness. Test the model inside the agent framework you will ship.

If you run Codex, GPT-6 Sol is the value pick. At XHigh it ties Astra's XHigh pass rate, 32.2%, for $318 against $898. Astra's Max buys 2.6 more points than Sol's Max for 2.6x the cost. On Near-Term, with each at its best effort, Sol reaches 50.7% for $81 against Astra's 52.2% for $283.

If you run Claude Code, move from Opus 5 to Opus 5.5. At Max effort, Opus 5.5 passes 7.3 points more for 0.90x the cost. Fable 5 at XHigh spends $4,340 for 48.7, and Opus 5.5 at Max scores 14.5 points higher for 31% of that, with the caveat that the efforts differ and the board flags the Fable 5 variant.

Do not default to maximum effort. Of the three models with a published effort ladder, only Astra does best at Max. Opus 5 does best at High, and GPT-6 Sol at XHigh.

Plan for failure on hard work. The best Overall configuration still fails 61.8% of tasks, every Codex configuration fails at least 65.8%, and the best Last-Exam result is six tasks of 38. The board gives no signal on which tasks will fail, so any production use needs a verification step that checks the deliverable before anyone relies on it. A frontier model on ALE is a strong drafter for professional files, not a finisher.

The LLM leaderboard compares models across benchmarks. If your work runs through CAD, simulation or other engineering software, the more useful number is a pass rate on your own deliverables; the engineering page covers how to set up that kind of evaluation in Klu.