What are OSWorld 2.0 and 2.1?
OSWorld 2.0 is a set of 108 long computer-use tasks. An AI agent gets a live Ubuntu virtual machine, sees it only through screenshots, and works it with mouse and keyboard actions until it decides the job is done. A grader then inspects the final state of the machine against a list of weighted checkpoints, 27.25 per task on average. OSWorld 2.1 is the same benchmark after a bug-fix release of the task files, and it is the release the maintainers now recommend.
XLANG Lab published the OSWorld 2.0 paper on June 28, 2026, with 36 authors led by Mengqi Yuan, Zilong Zhou and Xinzhuang Xiong. It takes over from the original OSWorld, the 2024 benchmark from Tianbao Xie et al. with 369 short, single-goal tasks. The 2.0 authors explain the rewrite in their first section: "Claude Opus 4.8 reaches 83.5% on OSWorld-Verified, suggesting desktop computer use is largely solved, the tasks behind this number are short and narrow, rarely spanning more than one or two applications ... High accuracy on such benchmarks therefore overstates real progress and obscures how rarely agents complete the end-to-end work that real deployment demands."
The change in scale is the whole idea. Per the paper:
- Median human operation time is about 1.6 hours per task. In OSWorld 1.0 it was about two minutes, so 2.0 tasks run about 48 times longer.
- OSWorld 1.0 trajectories average about 30 steps. The leading 2.0 agents average more than 300.
- A task touches 2.44 apps or services on average.
The best run in the paper, Claude Opus 4.8 at max thinking with batched tool calls, finishes 20.6% of tasks completely. By October 5, 2026, the official leaderboard listed Claude Opus 5 at 44.33% on release 2.1. Those two numbers come from different releases of the task files, and on this benchmark the release alone moves a score by double digits.
How the tasks and harness work
The 108 tasks span 7 professional domains and 21 subcategories. Research and Education plus Creative Production make up more than 40% of the set. Each task starts from a stateful user profile, not a blank machine. The virtual machine runs self-hosted websites for email, banking, team chat and business portals, next to desktop applications such as LibreOffice, GIMP, Thunderbird, VS Code, FreeCAD and REAPER. With 2.44 apps or services per task, the average task crosses more than two programs. Ubuntu is the only operating system the paper and the OSWorld-V2 repo document.
The agent sees screenshots and nothing else. How it acts depends on the vendor. In the paper's configuration, "Claude models act through the native claude_computer_use tool, while the remaining models emit pyautogui code actions." Batching differs too: "GPT-5.5 always batches its tool calls, whereas Claude models batch only when the batch tool is explicitly enabled."
Batching lets the agent send several actions in one step, and it is worth points. The same Opus 4.8 scores 20.6% strict pass batched and 18.5% with one call per step. Opus 4.7 goes from 13.9% single-call to 18.2% batched. The leaderboard keeps "batch tool" and "standard" rows apart, and a comparison across the two settings measures the tool setting as much as the model.
Step budget
Runs use a budget of 150, 300 or 500 steps. The paper's main results all use 500, which is also the leaderboard default, and Anthropic and Google both report at 500. The leaderboard carries the smaller budgets for two models, with the standard tool setting and max thinking.
| Model | 150 steps | 300 steps | 500 steps |
|---|---|---|---|
| Claude Opus 4.7 | 4.6% / 20.3% | 13.0% / 39.8% | 13.9% / 49.1% |
| Claude Sonnet 4.6 | 4.6% / 20.0% | 6.5% / 35.8% | 8.3% / 41.5% |
Scores are strict pass / partial score from the leaderboard data file. The 500-step rows match the paper's Table 3. Compare across columns within a row, not with batch-tool rows elsewhere on this page.
Opus 4.7 triples its strict pass rate between 150 and 500 steps, from 4.6% to 13.9%. Most of that gain arrives by 300 steps. From 300 to 500, strict pass moves 0.9 points while partial score climbs 9.3. The last 200 steps buy checkpoints, not finished tasks.
Effort, release and context handling
Two more settings each move a score further than the 6.6-point gap between the top two vendors on the leaderboard, and a third shifts Anthropic's own numbers.
- Reasoning effort. Opus 5 on release 2.1, full set, goes from 18.81% strict at low effort to 44.33% at max, a gain of 25.5 points. GPT-5.6 Sol on the v2026.08.08 offline set goes from 3.41% at "none" to 28.10% at max. This is test-time compute paying off on long tasks.
- Release. Opus 5 at max effort scores 31.43% strict on v2026.08.08 and 44.33% on v2.1, both on the full set. The bug-fix release adds 12.9 points.
- Context handling. Anthropic changed its harness between two system cards. The Fable 5.1 card limits how many screenshots the agent keeps. The Opus 5.5 card keeps every screenshot and compacts the context server-side above 100k tokens. Opus 5 scores 75.4% partial and 39.6% strict under the first harness, 74.0% and 37.2% under the second.
The paper reports no wall-clock timeout. Google states a 500-step limit. OpenAI has not published a step limit, harness, effort setting or run count for any OSWorld 2.0 score.
Access
Task classes sit in a gated Hugging Face dataset, xlangai/osworld_v2_tasks, and full assets in a second gated set, xlangai/osworld_v2_assets_gated. Per the README, the gating "reduces benchmark leakage and helps prevent evaluated agents from finding task answers, setup logic, or evaluator details online while executing a task." Provider images exist for Docker and AWS. No source states an instance type or a CPU or GPU requirement.
How scoring works
Every task has a list of weighted checkpoints, 27.25 on average, checked against the final machine state. OSWorld reports two numbers from them:
- Strict pass, which the leaderboard calls binary, counts a task only when every checkpoint passes.
- Partial score is the mean per-task checkpoint credit.
Most checkpoints are functional checks on machine state. Some use a model-graded rubric, a form of LLM-as-a-judge, and the authors cap those at 11.53% of the total score and at 50% of any single task. Anthropic's system cards name Claude Opus 4.8 as the grader for their runs.
Partial scores run far ahead of strict pass. Opus 5 at max effort on v2.1 earns 77.67% of checkpoint credit, yet leaves at least one checkpoint unmet on 55.67% of tasks. If you are deploying an agent to finish work unattended, strict pass is the number that answers "is the job done." I list both everywhere below and read strict pass first.
Full set and offline set
The repo lists an offline subset of 82 tasks that need no internet, and the leaderboard toggles between Full (108 tasks) and Offline. Google reports only the offline subset because it "has more robust verifiers." The two sets give different numbers for the same model and settings. Opus 5 at max effort on v2.1 scores 44.33% strict on the full set and 48.65% offline.
Versions
The README lists three active releases of the task files.
| Release | README status | Mocked websites | Who reports on it |
|---|---|---|---|
| v2026.06.24 | Active | No longer hosted at web.hku.icu | Every run in the paper |
| v2026.08.08 | Active | Hosted by the OSWorld team at site.hku.icu | Cross-vendor leaderboard rows, Google's Gemini 4 Argon protocol |
| osworld-v2.1 | Active (recommended), released 2026-09-16 as "a new bug-fix release" | Self-host only | Leaderboard Opus 5 rows, Simular's Sai agent |
Anthropic runs its own snapshots. The Fable 5.1 card used the August 2026 release plus fixes, and the Opus 5.5 card used "the September 10, 2026 versions" of the task files. Anthropic also reported setup and grading-script fixes back to the authors. The Opus 5.5 card calls its files OSWorld 2.0 and says its results "are not directly comparable to results reported on earlier releases of the benchmark or obtained with other harness configurations."
Current leaderboard
Every row in the official leaderboard data file, last updated October 5, 2026, carries the flag "official: true." The README reserves verified status for results the OSWorld team reproduces by running the agent code itself, or, for trusted institutions, checks from shared monitoring data and trajectories. The monitoring service returned an error on October 6, 2026, so every leaderboard row below is listed as official with verification not confirmed. No leaderboard row gives a run date. Scores read strict pass / partial score.
Cross-vendor, release v2026.08.08, offline set
This is the one table where two vendors run the same release, task set, step budget and tool setting.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5 (max) | Anthropic | 34.72% / 70.19% | v2026.08.08, offline 82 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. No cost listed |
| Claude Opus 5 (xhigh) | Anthropic | 33.38% / 70.14% | v2026.08.08, offline 82 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. No cost listed |
| Claude Opus 5 (high) | Anthropic | 32.69% / 65.92% | v2026.08.08, offline 82 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. No cost listed |
| GPT-5.6 Sol (max) | OpenAI | 28.10% / 64.13% | v2026.08.08, offline 82 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. Total run cost $1,116.72 |
| Claude Opus 5 (medium) | Anthropic | 27.13% / 60.26% | v2026.08.08, offline 82 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. No cost listed |
| GPT-5.6 Sol (xhigh) | OpenAI | 24.63% / 60.60% | v2026.08.08, offline 82 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. Total run cost $849.96 |
| Claude Opus 5 (low) | Anthropic | 24.61% / 55.22% | v2026.08.08, offline 82 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. No cost listed |
| GPT-5.6 Sol (high) | OpenAI | 20.00% / 56.03% | v2026.08.08, offline 82 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. Total run cost $679.32 |
| GPT-5.6 Sol (medium) | OpenAI | 17.56% / 49.01% | v2026.08.08, offline 82 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. Total run cost $379.08 |
| GPT-5.6 Sol (low) | OpenAI | 7.32% / 29.63% | v2026.08.08, offline 82 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. Total run cost $189.00 |
| GPT-5.6 Sol (none) | OpenAI | 3.41% / 16.43% | v2026.08.08, offline 82 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. Total run cost $96.12 |
Every row here compares with every other row, and with no other table on this page. Costs are the data file's total run cost; it does not say how many tasks the total covers, so no per-task cost is given.
At max effort, Opus 5 leads GPT-5.6 Sol by 6.6 strict points, 34.72% to 28.10%, and by 6.1 partial points. The gap holds down the effort ladder. Opus 5 at low effort, 24.61%, ties GPT-5.6 Sol at xhigh, 24.63%. Opus 5 at medium sits within a point of GPT-5.6 Sol at max.
The chart shows where effort stops paying. From none to low, GPT-5.6 Sol gains 13.2 partial points for $92.88 more in total run cost. Low to medium adds 19.4 points for $190.08. After medium the returns shrink: 7.0 points for $300.24 to reach high, 4.6 for $170.64 to xhigh, and 3.5 for $266.76 to max. Medium is the knee of the curve.
Cross-vendor, release v2026.08.08, full set
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5 (max) | Anthropic | 31.43% / 68.31% | v2026.08.08, full 108 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. No cost or tokens listed |
| Claude Opus 5 (xhigh) | Anthropic | 30.21% / 67.67% | v2026.08.08, full 108 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. No cost or tokens listed |
| Claude Opus 5 (high) | Anthropic | 28.98% / 63.79% | v2026.08.08, full 108 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. No cost or tokens listed |
| GPT-5.6 Sol (max) | OpenAI | 27.34% / 62.72% | v2026.08.08, full 108 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. No cost or tokens listed |
| Claude Opus 5 (medium) | Anthropic | 25.33% / 60.49% | v2026.08.08, full 108 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. No cost or tokens listed |
| Claude Opus 5 (low) | Anthropic | 22.32% / 54.89% | v2026.08.08, full 108 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. No cost or tokens listed |
Same release and setup as the offline table but all 108 tasks, so rows compare only within this table.
On the full set at max effort, Opus 5 leads GPT-5.6 Sol by 4.1 strict points and 5.6 partial points.
Release 2.1, full set
Claude Opus 5 is the only model with v2.1 rows on the official board.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5 (max) | Anthropic | 44.33% / 77.67% | v2.1, full 108 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. 107,919 output tokens per task. Offline 48.65% / 79.20% |
| Claude Opus 5 (high) | Anthropic | 36.89% / 68.92% | v2.1, full 108 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. 65,555 output tokens per task. Offline 39.74% / 70.55% |
| Claude Opus 5 (xhigh) | Anthropic | 33.33% / 73.11% | v2.1, full 108 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. 88,387 output tokens per task. Offline 34.67% / 72.03% |
| Claude Opus 5 (medium) | Anthropic | 33.01% / 64.85% | v2.1, full 108 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. 48,645 output tokens per task. Offline 37.18% / 66.25% |
| Claude Opus 5 (low) | Anthropic | 18.81% / 59.09% | v2.1, full 108 tasks, batch tool, 500 steps | Listed 2026-10-05 | Official leaderboard. 29,242 output tokens per task. Offline 22.08% / 60.34% |
One model on the bug-fixed task files. These rows do not compare with v2026.08.08 rows, which grade against older files. No cost is listed.
Strict pass drops from high to xhigh, 36.89% to 33.33%, while partial score rises from 68.92% to 73.11%. The offline set shows the same dip. Max is the only setting that pulls strict pass well clear of high, 7.4 points higher, and it spends 3.7 times the output tokens of low effort to get there.
Third-party agent, release 2.1
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Sai with Claude Opus 5 (max) | Simular | 41.67% / 79.38% | Simular's Sai agent, v2.1, full 108 tasks, 500 steps | Listed 2026-10-05 | Run and submitted by Simular. $14.34 per task. Trajectories |
An agent built around Opus 5, not a bare model. It compares only with Opus 5 at max effort in the release 2.1 full-set table.
Against bare Opus 5 at max on the same release and set, Sai is 2.7 strict points lower and 1.7 partial points higher. The extra scaffold earns more checkpoint credit and finishes fewer tasks.
Anthropic system cards
The Claude Opus 5.5 System Card, dated September 22, 2026, re-ran three models under one protocol.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 48.7% / 81.8% | 2026-09-10 task files, full set, 500 steps, max effort, pass@1 averaged over 5 runs, Opus 4.8 grader, all screenshots kept with 100k-token compaction | 2026-09-22 | Opus 5.5 card, section 8.13.3. Vendor-reported |
| Claude Fable 5.1 | Anthropic | 42.8% / 80.7% | Same as above | 2026-09-22 | Opus 5.5 card, section 8.13.3. Vendor-reported |
| Claude Opus 5 | Anthropic | 37.2% / 74.0% | Same as above | 2026-09-22 | Opus 5.5 card, section 8.13.3. Vendor-reported |
Anthropic-run on Anthropic's harness and task snapshot. Rows compare with each other only.
The Claude Fable 5.1 and Mythos 5.1 System Card, dated September 1, 2026, used a different harness.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Fable 5.1 / Mythos 5.1 | Anthropic | 41.7% / 77.9% | August 2026 release plus fixes, 500 steps, max effort, pass@1 over 5 runs, Opus 4.8 grader, screenshot-count-limited harness | 2026-09-01 | Fable 5.1 card, section 8.14.3. Vendor-reported. Safeguard interventions scored zero |
| Claude Opus 5 | Anthropic | 39.6% / 75.4% | Same as above | 2026-09-01 | Fable 5.1 card, section 8.14.3. Vendor-reported. Re-run |
| Claude Fable 5 / Mythos 5 | Anthropic | 36.1% / 72.9% | Same as above | 2026-09-01 | Fable 5.1 card, section 8.14.3. Vendor-reported. Re-run, safeguard interventions scored zero |
A different task snapshot and context harness from the Opus 5.5 card, so rows compare only with each other.
In Anthropic's newest run, Opus 5.5 leads Fable 5.1 by 5.9 strict points and Opus 5 by 11.5, while partial scores sit 1.1 and 7.8 points apart. Opus 5.5 finishes more tasks outright, not just more checkpoints. Its 48.7% is the highest strict figure on this page, but it comes from Anthropic's own run, and the card's summary table lists no GPT-5.6 Sol score. In the Fable card, production safeguards were on, and Anthropic scored Fable 5.1 and Fable 5 at zero on tasks where the safeguards intervened.
Google and OpenAI
Google evaluates Gemini 4 Argon on the offline subset with the max of 3 runs, one attempt each, its Gemini CUA harness with parallel batch tool calls, compaction per the Opus 5.5 card, pyautogui actions, screenshot-only observation, 500 steps, the official evaluator, the 08.08 patch, safety filters on, and the highest thinking setting, per its evaluation methodology. Google reports no Claude row, saying Anthropic "only reports aggregated results for the online and offline subsets combined." The PDF's results table is an image with no text layer, so this page carries no Argon score. OpenAI's GPT-6 Astra post was not readable, and Astra is not on the official board, which lists GPT-5.6 Sol.
Benchmark authors' runs, release 2026.06.24
The paper's Table 3 covers seven models, all run by the authors at 500 steps on the 2026.06.24 release. The leaderboard data file carries the same scores. Costs are the paper's approximate figures per task. The tables split by tool mode and thinking setting, because the paper changed both between models.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 4.8 | Anthropic | 20.6% / 54.8% | Batch tool, max thinking, 500 steps | 2026-06-28 | Paper Table 3. About $72.4 per task, 224K output tokens |
| Claude Opus 4.7 | Anthropic | 18.2% / 48.9% | Batch tool, max thinking, 500 steps | 2026-06-28 | Paper Table 3. About $33.6 per task, 150K output tokens |
Batched mode at max thinking; compares within this table only.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| GPT-5.5 | OpenAI | 13.0% / 49.5% | Batched, xhigh effort, 500 steps | 2026-06-28 | Paper Table 3. About $25.5 per task, 37.1K output tokens |
Batched, but at xhigh effort where the Claude rows ran at max. The authors published no same-effort GPT-5.5 row, so this row does not rank against the Claude table above.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Claude Opus 4.8 | Anthropic | 18.5% / 49.3% | Single tool call, max thinking, 500 steps | 2026-06-28 | Paper Table 3. About $76.1 per task, 259.5K output tokens |
| Claude Opus 4.7 | Anthropic | 13.9% / 49.1% | Single tool call, max thinking, 500 steps | 2026-06-28 | Paper Table 3. About $35.8 per task, 150.5K output tokens |
| Claude Sonnet 4.6 | Anthropic | 8.3% / 41.5% | Single tool call, max thinking, 500 steps | 2026-06-28 | Paper Table 3. About $22.3 per task, 185.9K output tokens |
Single tool call per step at max thinking; compares within this table only.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| MiniMax M3 | Not listed | 4.6% / 22.3% | Single tool call, thinking on, 500 steps | 2026-06-28 | Paper Table 3. About $2.4 per task, 70.8K output tokens |
| Kimi 2.6 | Not listed | 4.6% / 22.1% | Single tool call, thinking on, 500 steps | 2026-06-28 | Paper Table 3. About $6.6 per task, 63.0K output tokens |
Open-weight models with thinking simply switched on; compares within this table only.
| Model | Organization | Score | Harness and setup | Date | Source and notes |
|---|---|---|---|---|---|
| Qwen 3.7-Plus | Not listed | 2.8% / 21.5% | Single tool call, thinking off, 500 steps | 2026-06-28 | Paper Table 3. About $3.8 per task, 28.9K output tokens |
Thinking off, so it ranks against none of the other paper tables.
Batched Opus 4.8 leads batched Opus 4.7 by 2.4 strict points at more than twice the cost per task. In single-call mode Opus 4.8 leads Opus 4.7 by 4.6 strict points and only 0.2 partial points, so the newer model's edge shows up in finished tasks. The best open-weight row, MiniMax M3 with thinking on, reaches 22.3% partial against 49.3% for single-call Opus 4.8 at max thinking, at about a thirtieth of the cost per task, under a different thinking setting.
Historical results
| Date | Release | Model | Score | Who ran it | Notes |
|---|---|---|---|---|---|
| 2024-04-11 | OSWorld 1.0, 369 tasks | Best model in the paper | 12.24% success | Benchmark authors | Humans scored 72.36%. OSWorld paper |
| 2024-10 | OSWorld 1.0 | Claude 3.5 Sonnet | 14.9% screenshot-only, 22.0% with more steps | Anthropic | Next-best system 7.8% screenshot-only. Anthropic |
| 2026-06-28 | 2.0, 2026.06.24 release | Claude Opus 4.8 | 20.6% / 54.8% | Benchmark authors | Batch tool, max thinking |
| 2026-09-01 | 2.0, August 2026 release plus fixes | Claude Fable 5.1 | 41.7% / 77.9% | Anthropic | Screenshot-count-limited harness |
| 2026-09-22 | 2.0, 2026-09-10 task files | Claude Opus 5.5 | 48.7% / 81.8% | Anthropic | All screenshots kept, 100k-token compaction |
| Listed 2026-10-05 | 2.0, v2026.08.08 | Claude Opus 5 (max) | 31.43% / 68.31% | Leaderboard listing | Full set. GPT-5.6 Sol at max 27.34% / 62.72% |
| Listed 2026-10-05 | 2.1 | Claude Opus 5 (max) | 44.33% / 77.67% | Leaderboard listing | Full set |
OSWorld 1.0 is a different task set and compares with no 2.0 score. Each 2.0 row uses a different task release or harness, so the rows date the protocol and do not form a score trend.
Where agents fail
The paper names five recurring failures.
- Information grounding. Agents drop constraints, miss updates that arrive mid-task, and act without evidence.
- Perception-action timing. On moving pop-ups and streams, the agent acts on a screenshot that is already stale.
- Domain artifact interpretation. Agents misread FreeCAD geometry and video timelines.
- Verification gaps. An agent sees a problem and does not revise its plan, or skips the final check.
- State drift. Information held only in compressed chain of thought gets lost over a long run.
Completion collapses as tasks get longer. In the authors' runs of Opus 4.8, Opus 4.7 and GPT-5.5 on the 2026.06.24 release, binary completion is 20 to 24% for tasks under 45 minutes of human time. No model exceeds 10% in the 137 to 163 minute bin, and every model scores 0% above 163 minutes. The median task takes a human 1.6 hours, well past the 45-minute mark where the best completion rates stop.
Agents also do not spend effort on fixing their own mistakes. Correction, meaning recovery plus repair, stays below 7% of agent effort in every model the paper measured. The budget goes to visual grounding (15.5%), tool semantics (13.8%) and information extraction (12.8%), with 9.8% on verification. Those shares line up with the verification gap above and with the distance between partial and strict scores everywhere on this page.
The models solve tasks in different styles. GPT-5.5 works programmatically, using code or APIs 78% of the time. Opus 4.7 stays in the UI and loses convergence on exact details.
Then there is safety. Across 216 trajectories, GPT-5.5 and Opus 4.7 on all 108 tasks, 14% of tasks involved extracting hidden app state and 33% involved bypassing user-visible interfaces. The paper's examples include credential leaks, resource exhaustion and working around the UI when blocked. Before you rely on an OSWorld score, put the trajectories behind it through a red-teaming review.
Measurement problems
- Contamination. Gating the task files is the only control either source describes. A search of the paper and README for leakage, contamination, memorization and decontamination turns up no training-data overlap measurement. Task creators drew ideas from public tutorials, Reddit threads and documentation, all public-web material, and the paper's limitations section warns that "agents may also learn to exploit benchmark-specific artifacts in self-hosted environments over time."
- Task fixes. Anthropic reported setup and grading-script fixes to the authors, and the 2.1 bug-fix release lifted Opus 5 at max from 31.43% to 44.33% strict on the full set.
- Protocol. Opus 5 at max effort scores 70.19% partial offline on v2026.08.08, 79.20% offline on v2.1, 74.0% in the Opus 5.5 card and 75.4% in the Fable 5.1 card. Same model, four numbers.
- Safety filters. Google returns flagged responses as empty strings. Anthropic scored zero where Fable 5.1's safeguards intervened.
- Run counts. Google takes the max of 3 runs. Anthropic averages 5. A max of 3 runs sits at or above the average of those same runs by construction.
How it compares to related benchmarks
OSWorld 1.0 and OSWorld-Verified share the environment family with 369 short tasks, about two minutes of human time and about 30 agent steps each. Claude Opus 4.8 reached 83.5% on OSWorld-Verified, the figure that prompted the 2.0 rewrite. OSWorld 2.0 keeps the live virtual machine and swaps in 108 tasks that take a human 1.6 hours, graded by checkpoints.
Agent's Last Exam is the closest neighbor. It tests economically valuable professional work in real software, in a Linux or Windows sandbox, and grades the deliverable the agent returns, with deterministic judges on 93.2% of its open-sourced workflows (paper). Its leaderboard score is a partial-credit mean. OSWorld grades the final state of the whole machine and also reports strict pass, so it tells you how often the agent finished, not only how much it got right.
ScreenSpot-Pro isolates one skill OSWorld depends on. Each item is a single screenshot of a professional application and one predicted click point, with no environment and no episode. It measures GUI grounding. OSWorld measures whether grounding, planning and verification hold together over hundreds of steps.
AutomationBench covers cross-application business workflows too, but over simulated SaaS apps and API endpoints through Search and Execute tools, scored binary. OSWorld agents drive a GUI from screenshots. If your agent will call APIs, AutomationBench is the closer test. If it has to work software that has no API, OSWorld is.
For code and terminal agents, see SWE-bench and Terminal-Bench 4.0. GAIA is another agentic benchmark worth knowing. The wider map is in LLM benchmarks and LLM evaluation.
What OSWorld 2.0 means for teams choosing a model
Claude Opus 5 leads on the only like-for-like table. On v2026.08.08 at max effort, Opus 5 beats GPT-5.6 Sol by 6.6 strict and 6.1 partial points on the offline set, and by 4.1 strict and 5.6 partial points on the full set. On the offline set, Opus 5 at low effort matches GPT-5.6 Sol at xhigh. Opus 5 is also the only model with rows on the recommended 2.1 release. Among the frontier models a team would shortlist for desktop automation today, that makes Opus 5 the model with the strongest public evidence.
Newer models have no comparable score. Neither GPT-6 Astra nor Gemini 4 Argon has a readable OSWorld 2.0 number, and the newest OpenAI model on the board is GPT-5.6 Sol. Anthropic's card puts Opus 5.5 at 48.7% strict, 11.5 points above Opus 5 on the same harness, but no other vendor ran that harness. If you are weighing Opus 5.5 against Astra or Argon for computer use, OSWorld cannot settle it yet.
Pay for effort up to medium, then decide. GPT-5.6 Sol gains 32.6 partial points from none to medium for $282.96 in added total run cost. Going from medium to max adds 15.1 points for 2.9 times the cost. For Opus 5 on v2.1, max beats high by 7.4 strict points, while xhigh scores 3.6 strict points below high. If you run Opus 5 above high, run it at max.
Read strict pass before partial. Opus 5.5 scores 81.8% partial and 48.7% strict in Anthropic's run, so 51.3% of tasks fail at least one checkpoint. The paper's explanation is that agents spend under 7% of effort on detecting and repairing their own errors and skip the final checks completion depends on. Long tasks make it worse. In the authors' runs, completion was 20 to 24% for tasks under 45 minutes of human time and zero above 163 minutes. Budget for human review of anything longer than an hour of human work.
Use 2.1 numbers, and check the release before comparing. The bug-fix release added 12.9 strict points to Opus 5 at max. A vendor score on an older release or a private snapshot is a different measurement.
The Klu LLM leaderboard tracks these models across other benchmarks. To score agents against checkpoints on your own back-office workflows, the operations page shows how teams set up the same kind of eval in Klu.