ScreenSpot-Pro

A GUI grounding benchmark that tests whether a model can find and click the one correct element on full-resolution screenshots of professional desktop software

Computer useAccuracy, greedy decoding, micro-average

Stephen M. Walker II · Co-Founder / CEO · October 6, 2026 · 11 min read

What is ScreenSpot-Pro?

ScreenSpot-Pro is a GUI grounding benchmark, a narrow computer vision test for screen agents. The model gets a full-screen screenshot of a professional desktop application and a short natural-language instruction. It has to return the location of the one UI element that satisfies the instruction. Nothing else is asked. There is no plan to write, no task to finish, no environment to drive. The benchmark checks whether the model can find a small icon on a 5K screen and point at it.

Kaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo, Yuchen Tian, Jing Ma, Zhiyong Huang and Tat-Seng Chua built it across the National University of Singapore, East China Normal University and Hong Kong Baptist University. They released the dataset and paper on 2025-01-04 and posted it to arXiv on 2025-04-04.

The reason it exists is that the earlier ScreenSpot benchmark had become too easy to be useful. ScreenSpot cropped screenshots down to local regions of consumer mobile, web and light desktop apps, and the average target covered 2.01% of the image. ScreenSpot-Pro keeps the whole screen and moves to IDEs, CAD tools, creative suites, scientific packages, office software and operating system settings. The average target now covers 0.07% of the screenshot, about one twenty-ninth of the old size.

That change broke the models of the day. At release the best of ten models tested, OS-Atlas-7B, scored 18.9%. GPT-4o scored 0.8%. By October 2026 the published scores sit in the 70s and 80s, and the highest of them all wrap the model in cropping, zooming or multi-step search. Anthropic, Google DeepMind, Qwen, H Company and OpenAI all cite it among their LLM benchmarks, and the gap between a model's raw score and its wrapped score is the most useful thing it measures.

How the task works

The dataset holds 1,581 instructions, each on its own screenshot. They come from 23 applications across five industries and three operating systems: Windows, macOS and Linux. The paper groups the applications into six genres.

  • Development and programming. VS Code, PyCharm, Android Studio, Quartus, VMware Fusion.
  • Creative. Photoshop, Premiere, Illustrator, Blender, FruitLoops Studio, Unreal Engine, DaVinci Resolve.
  • CAD and engineering. AutoCAD, SolidWorks, Inventor, Vivado.
  • Scientific and analytical. MATLAB, Origin, Stata, EViews.
  • Office suite. Word, PowerPoint, Excel.
  • Operating system commons. Windows 11, macOS Sonoma, Ubuntu 24.04.

Annotators were experts with at least five years in each application. They recorded their normal workflow with a silent screenshot tool on monitors above 1080p with display scaling off, then drew the target box themselves. Resolutions in the set include 2560x1440, 3840x2160, 5120x2880, 2560x1664 and 2880x1800. At least two annotators reviewed every instance. Ambiguous instructions were rewritten until exactly one target existed, and boxes cover only the clickable region of the element.

Targets split 62.6% text and 37.4% icon. An element counts as an icon only when it has no text label, so a labeled toolbar button counts as text. The icon share is where models fail, as the failure modes section shows.

Inputs, output and scoring

The input is the screenshot plus the instruction. The output is a single click coordinate. Models that emit a bounding box have the center of that box scored as the click. A prediction is correct when the point lands inside the ground-truth box. The model gets no accessibility tree and no DOM, only pixels.

The headline number is accuracy over all 1,581 instructions. The official grounding leaderboard reports the micro-average under greedy decoding, with breakdowns by application, by genre and by text versus icon. The maintainers collect submissions directly rather than running models themselves.

Each item is static. One screenshot in, one prediction out, with no episode, no sandbox, no step limit and no time limit. The reference evaluation scripts are on GitHub under the MIT license, and the dataset is on Hugging Face. A full run is 1,581 model calls without a wrapper, small enough to rerun on every model release.

How the labs run it

"One screenshot, one click" is the benchmark's definition. Frontier labs rarely report it that way. Each lab that publishes a number uses its own setup, and most add a tool or a loop around the model.

  • Anthropic reports Claude in the Opus 4.8 System Card (section 8.12.5) with adaptive thinking at maximum effort, in two modes. One gives the model no tools. The other gives it Python tools. Each score is a five-run average.
  • Google DeepMind runs Gemini 3 Pro with a "capture screenshot" function-calling tool that passes the screenshot back to the model, at media_resolution "extra_high", per its evaluation methodology PDF. DeepMind ran the non-Gemini models itself through provider APIs because no self-reported or leaderboard numbers existed.
  • Qwen gives Qwen3-VL a computer_use tool limited to left_click and mouse_move, tells it to center the cursor tip on the target, and lets it give up if it cannot find the element (Qwen3-VL Technical Report, Appendix B.11).
  • H Company reports Holo2 both single-step and with "agentic localization," which refines the guess for up to three steps over the 4K screenshot (Holo2 release post).
  • OpenAI says GPT-6 Astra "achieves state-of-the-art results on Agents' Last Exam, AutomationBench, and ScreenSpot Pro" in its announcement. It has published no score, effort setting, scaffold or resolution. The GPT-6 Astra System Card and the GPT-6.1 Sol addendum have no ScreenSpot-Pro section.

Image resolution also changes scores. Anthropic's Opus 4.7 release raised image input to 2,576 pixels on the long edge, about 3.75 megapixels and more than three times the limit of earlier Claude models. ScreenSpot-Pro screenshots run up to 5120x2880, so even that limit means a 5K screen is downscaled before the model sees it.

Versions

ScreenSpot-Pro has no numbered versions. The only dated maintainer updates are in the repo README.

  • 2025-01-04. Paper and dataset released.
  • 2025-02-21. The README lists projects already using it as a benchmark, including OmniParser v2, Qwen2.5-VL, UI-TARS, UGround and AGUVIS.
  • 2025-05-19. SE-GUI results added, 47.2% for the 7B model and 35.9% for the 3B, trained on 3,000 open-source samples.

Every task also has a Chinese instruction, translated with GPT-4 and then reviewed. That set ships as ScreenSpot-Pro-CN, described in paper Appendix A and run with the repo script run_ss_pro_cn.sh. The leaderboard page was last updated 2026-09-10.

Current leaderboard

Every score on this page is vendor-reported or leaderboard-submitted. No independent run of a frontier model on ScreenSpot-Pro exists, and OpenAI has published no number.

The highest published score is Claude Opus 4.8 at 87.9% with Python tools, under Anthropic's harness. The top of the official leaderboard is GroundCascade at 85.1%, a cascade over two open models. Those two numbers come from different harnesses and do not rank against each other. The tables below are grouped so that rows inside each table do compare.

Anthropic harness

One harness, so the four rows compare directly. Do not compare them to the leaderboard or DeepMind tables.

ModelOrganizationScoreHarness and setupDateSource and notes
Claude Opus 4.8Anthropic87.9%Python tools, adaptive thinking, max effort, five-run average2026-05-28Opus 4.8 System Card, section 8.12.5. Tools add 5.6 points.
Claude Opus 4.7Anthropic87.6%Python tools, adaptive thinking, max effort, five-run averageModel 2026-04-16, reported 2026-05-28Same card. Tools add 8.1 points.
Claude Opus 4.8Anthropic82.3%No tools, adaptive thinking, max effort, five-run average2026-05-28Same card. 95% confidence intervals appear only in a figure, with no numbers in the text.
Claude Opus 4.7Anthropic79.5%No tools, adaptive thinking, max effort, five-run averageModel 2026-04-16, reported 2026-05-28Same card. The card's GPT-5.5 and Gemini 3.1 Pro columns show "-" for this benchmark.
AnthropicAdaptive thinking, max effort, five-run averageAccuracywww-cdn.anthropic.com

Opus 4.8 beats Opus 4.7 by 2.8 points without tools and by 0.3 points with them. The tools close almost the entire gap between the two models.

Google DeepMind's Gemini 3 Pro table

DeepMind ran every row. The Gemini rows compare with each other. The non-Gemini rows use settings DeepMind did not document, so they work only as a rough reference to older models, not as a comparison with the Anthropic table.

ModelOrganizationScoreHarness and setupDateSource and notes
Gemini 3 ProGoogle72.7%Screenshot function-calling tool, media_resolution "extra_high"2025-11Gemini 3 Pro evaluation PDF, page 3. Model id gemini-3-pro-preview, pass@1.
Gemini 3 ProGoogle60.5%Same tool, media_resolution "high"2025-11Same PDF, from the text. One setting is worth 12.2 points.
Claude Sonnet 4.5Anthropic36.2%Run by DeepMind through the official provider API2025-11Same PDF. DeepMind documents the screenshot tool only for Gemini 3.
Gemini 2.5 ProGoogle11.4%Run by DeepMind through the official provider API2025-11Same PDF.
GPT-5.1OpenAI3.5%Run by DeepMind through the official provider API2025-11Same PDF.
Google DeepMindAccuracystorage.googleapis.com

The PDF says of extra_high, "This setting is coming soon to our API." The 72.7% figure used a setting customers could not yet select when it was published.

Official leaderboard, wrapped entries

Greedy decoding, micro-average, leaderboard-submitted. Each entry adds a submitter-built wrapper over open models. These rows compare with the single-pass leaderboard table only as the wrapper-versus-base pairs called out below.

ModelOrganizationScoreHarness and setupDateSource and notes
GroundCascade (Indeed-UI-32B + KV-Ground-8B)Not listed85.1%Training-free geometric-consistency cascade, fixed 0.5 zoom, 2.574 model calls per exampleListed as of 2026-09-10Leaderboard. Top entry.
GroundCascade (Indeed-UI-32B + Indeed-UI-8B)Not listed85.0%Same cascade, 2.572 model calls per exampleListed as of 2026-09-10Leaderboard.
Indeed-UI-32B-zoominNot listed82.7%Zoom-inListed as of 2026-09-10Leaderboard. Base weights score 73.3%, so zoom adds 9.4 points.
Indeed-UI-8B-zoom-inNot listed81.0%Zoom-inListed as of 2026-09-10Leaderboard. Base 73.4%, so plus 7.6.
KV-Ground-8B + Qwen3.5-27B Consistency RouterNot listed80.9%Two-model ensembleListed as of 2026-09-10Leaderboard.
KV-Ground-GuiOwl1.5-0315-8B-ZoomInNot listed80.5%Zoom-inListed as of 2026-09-10Leaderboard. Base 73.2%, so plus 7.3.
Holo2-235B-A22B (Agentic)H Company78.5%Agentic localization, up to three refinement steps2026-02-03Holo2 post and leaderboard. Base 70.6%, so plus 7.9.
MAI-UI-32B (MVP)Not listed77.5%Multi-view prediction clustering, "MVP on attention layer 48; max inference process is 2"Listed as of 2026-09-10Leaderboard. Base 67.9%, so plus 9.6.
MAI-UI-32B (Zoom In)Not listed73.5%Zoom-inListed as of 2026-09-10Leaderboard. Base 67.9%, so plus 5.6.
ScreenSpot-Pro official leaderboardAccuracy, greedy decoding, micro-averagegui-agent.github.io

The leaderboard JSON holds 99 entries. None of the entry names contains Opus, Sonnet, Gemini or GPT-6. The top of the official board is open 8B to 32B models with wrappers, not frontier APIs.

Official leaderboard, single-pass entries

Same protocol, one forward pass per example, no wrapper. Legacy rows predate current models and used their own prompts, so the bottom of this table is history more than competition.

ModelOrganizationScoreHarness and setupDateSource and notes
Indeed-UI-8BNot listed73.4%Single passListed as of 2026-09-10Leaderboard. Top single pass.
Indeed-UI-32BNot listed73.3%Single passListed as of 2026-09-10Leaderboard. The 8B model edges the 32B by 0.1 points.
KV-Ground-GuiOwl1.5-0315-8BNot listed73.2%Single passListed as of 2026-09-10Leaderboard.
Holo2-235B-A22BH Company70.6%Single step2026-02-03Holo2 post and leaderboard.
MAI-UI-32BNot listed67.9%Single passListed as of 2026-09-10Leaderboard.
Holo1.5-72BNot listed63.3%Single passListed as of 2026-09-10Leaderboard.
UI-TARS-1.5Not listed61.6%Single passListed as of 2026-09-10Leaderboard.
Seed-1.5-VLNot listed60.9%Single passListed as of 2026-09-10Leaderboard.
Qwen2.5-VL-72B-InstructQwen53.3%Single pass2025 entryLeaderboard.
UI-TARS-72BNot listed38.1%Single pass2025 entryLeaderboard.
OS-Atlas-7BNot listed18.9%Single pass2025-01-04Best model at release, per the paper.
GPT5-minimal (resized)OpenAI18.5%Undocumented, image resizedNot listedEntry has no description or link. Submitter unknown.
Claude (Computer Use)Anthropic17.1%UndocumentedNot listedEntry has no description or link. Submitter unknown.
GPT5-high (resized)OpenAI6.0%Undocumented, image resizedNot listedEntry has no description or link. Submitter unknown. Higher reasoning scored 12.5 points lower.
GPT-4oOpenAI0.8%Single pass2025-01-04Paper baseline.
ScreenSpot-Pro official leaderboardSingle passAccuracy, greedy decoding, micro-averagegui-agent.github.io

The three frontier-API rows near the bottom tell you nothing about current Claude or GPT models. Nobody knows who ran them, when, or with what prompt.

Holo2, single step versus agentic localization

One harness from H Company, so rows compare directly. All eight rows also appear on the official leaderboard.

ModelOrganizationScoreHarness and setupDateSource and notes
Holo2-235B-A22BH Company78.5%Agentic localization, up to three steps2026-02-03Holo2 post. Plus 7.9 over single step.
Holo2-30B-A3BH Company75.2%Agentic localization, up to three stepsListed as of 2026-09-10Leaderboard entry "Holo2-30B-A3B (Agentic)". Plus 9.1.
Holo2-8BH Company71.4%Agentic localization, up to three stepsListed as of 2026-09-10Leaderboard entry "Holo2-8B (Agentic)". Plus 12.5.
Holo2-235B-A22BH Company70.6%Single step2026-02-03Holo2 post.
Holo2-4BH Company68.6%Agentic localization, up to three stepsListed as of 2026-09-10Leaderboard entry "Holo2-4B (Agentic)". Plus 11.4.
Holo2-30B-A3BH Company66.1%Single stepListed as of 2026-09-10Leaderboard entry "Holo2-30B-A3B".
Holo2-8BH Company58.9%Single stepListed as of 2026-09-10Leaderboard entry "Holo2-8B".
Holo2-4BH Company57.2%Single stepListed as of 2026-09-10Leaderboard entry "Holo2-4B".
H CompanyAccuracyhcompany.ai

The 8B model with three steps of refinement scores 71.4%, ahead of the 235B model in a single step at 70.6%. The 8B and 4B models gain 12.5 and 11.4 points from the loop. The 235B gains 7.9.

Qwen3-VL

Qwen's own harness, with the click-only computer_use tool. It does not compare with the other tables, and Qwen's own Table 2 shows "-" for Gemini 2.5 Pro, GPT-5 and Claude Opus 4.1.

ModelOrganizationScoreHarness and setupDateSource and notes
Qwen3-VL-235B-A22BQwen62.0% Instruct, 61.8% Thinkingcomputer_use tool, left_click and mouse_move only2025-12-01Qwen3-VL Technical Report, Tables 2 and 3.
Qwen3-VL-30B-A3BQwen60.5% Instruct, 57.3% ThinkingSame2025-12-01Same report.
Qwen3-VL-32BQwen57.9% Instruct, 57.1% ThinkingSame2025-12-01Same report.
Qwencomputer_use tool, left_click and mouse_move onlyAccuracy, Instruct and Thinkingarxiv.org

The Thinking variants score at or below Instruct at every size, by 0.2 to 3.2 points. Under Qwen's harness, extra reasoning bought nothing on grounding.

Historical progression

Each row carries its own harness. Read the column as a timeline of what was published, not as one model improving on one setup.

DateModel or methodScoreHarness and source
2025-01-04OS-Atlas-7B18.9%Single pass, best of ten models in the paper. UGround-7B 16.5%, AriaUI 11.3%, Qwen2-VL-7B 1.6%, SeeClick 1.1%.
2025-01-04GPT-4o0.8%Single pass, paper Tables 2 and 3.
2025-04-04ScreenSeekeR (GPT-4o + OS-Atlas-7B)48.1%Training-free search, paper Table 4. ReGround 40.2%, Iterative Narrowing 31.9%, Iterative Focusing 31.0%.
2025-05-19SE-GUI-7B47.2%Repo README. SE-GUI-3B 35.9%.
2025, no per-entry datesQwen2.5-VL-72B-Instruct53.3%Leaderboard. Also UI-TARS-72B 38.1%, UI-TARS-7B 35.7%, Qwen2.5-VL-7B-Instruct 26.8%, GUI-Actor-2.5VL-7B 44.6%.
2025-11Gemini 3 Pro72.7%Screenshot tool, extra_high resolution, DeepMind PDF.
2025-12-01Qwen3-VL-235B-A22B Instruct62.0%Qwen computer_use tool, Qwen3-VL report.
2026-02-03Holo2-235B-A22B78.5%Agentic localization, 70.6% single step, Holo2 post.
2026-04-16 to 2026-05-28Claude Opus 4.7, then Opus 4.887.6%, 87.9%Anthropic harness with Python tools. 79.5% and 82.3% without tools.
2026-09-03GPT-6 AstraNot publishedOpenAI claims state-of-the-art with no number.
2026-09-10GroundCascade85.1%Leaderboard top entry, cascade over two open models.

The most telling row is from April 2025. GPT-4o scored 0.8% when asked to click directly. Used as the planner in ScreenSeekeR, telling OS-Atlas-7B where to look, it lifted that model from 18.9% to 48.1%. A model that cannot click can still search well.

Failure modes and limitations

Icons. At release OS-Atlas-7B scored 28.1% on text targets and 4.0% on icons. The paper blames professional toolbars with many functions, apps that assume the user already knows them, and icons with domain meanings that rarely show up in web training data.

Target size. The average target covers 0.07% of the screen. The paper's Figure 2 shows accuracy falling as box area shrinks for SeeClick, OS-Atlas, UGround and Qwen2-VL on ScreenSpot-v2. ScreenSpot-Pro's average target is about one twenty-ninth the size of ScreenSpot's.

Resolution and crop settings outweigh model choice. Gemini 3 Pro moves 12.2 points between the "high" and "extra_high" media resolution settings. In the paper's Table 5, ReGround with OS-Atlas-7B scores 25.1% at 512x512 crops, 34.2% at 768, 40.2% at 1024 and 40.1% at 1280, while UGround peaks at 768 with 28.8% and falls to 26.3% at 1280. There is no single best crop size across models. On the leaderboard, GPT5-high (resized) scored 6.0% against 18.5% for GPT5-minimal (resized).

Every wrapper helps. Every model reported both with and without a wrapper gains from it. Anthropic's Python tools add 5.6 points to Opus 4.8 and 8.1 to Opus 4.7. Leaderboard wrappers add 5.6 to 9.6 points over the same weights. Holo2's agentic loop adds 7.9 to 12.5. A score without a stated wrapper is not usable for comparison.

Grounding is not task completion. The paper's Limitations section says the benchmark leaves out agent planning and execution of the kind OSWorld tests, partly to avoid legal risk from software licensing. A model that clicks the right button still has to know which button comes next.

Small per-app counts. With 1,581 items across 23 apps, per-app scores rest on thin samples. Stata has 49 targets and EViews has 50. Read per-application breakdowns as directional.

Contamination is unchecked. The dataset has been public on Hugging Face and GitHub since 2025-01-04. None of the Anthropic system card, the DeepMind methodology, the Qwen3-VL report or the Holo2 post reports a contamination or decontamination check for ScreenSpot-Pro.

Reporting gaps. Anthropic gives its 95% confidence intervals only in a figure. OpenAI has published no number, effort, scaffold or run count. DeepMind documents its screenshot tool only for Gemini 3. No lab publishes cost or latency for the task, which is why this page has no cost or time chart.

ScreenSpot is the predecessor, with cropped screenshots from mobile, web and light desktop apps. Its average target is 2.01% of the image against 0.07% here. Treat ScreenSpot scores as a separate benchmark, not an easier subset.

OSWorld runs 361 tasks end to end with an agent in a live Ubuntu VM. ScreenSpot-Pro isolates one step of that loop, putting the cursor on the right pixel, and drops planning and execution. Anthropic reports Opus 4.8 at 83.4% on OSWorld-Verified in the same system card. If your agent fails tasks, ScreenSpot-Pro tells you whether perception is the cause. OSWorld tells you whether the whole agent works.

MMMU tests image understanding with text answers and no pixel output. Answering a question about an image in words is a different skill from returning the coordinate of a toolbar button, so a high MMMU score does not predict ScreenSpot-Pro. When you shortlist large multimodal models for a screen agent, look at ScreenSpot-Pro instead.

AutomationBench measures enterprise workflow automation through apps and APIs with no screenshot grounding at all. OpenAI's Astra claim pairs ScreenSpot-Pro with AutomationBench and Agents' Last Exam.

What ScreenSpot-Pro means for teams choosing a model

Each point below stays within one harness.

  1. Turn on a crop or zoom tool before comparing models. Under Anthropic's harness, Python tools take Opus 4.8 from 82.3% to 87.9% and Opus 4.7 from 79.5% to 87.6%. Comparing Claude versions without tools measures a setup you should not ship.
  2. Do not upgrade for grounding alone if you already crop. Opus 4.8 gains 2.8 points over Opus 4.7 without tools and 0.3 points with them. If your agent already zooms, a 0.3-point gain will not show up in your traffic.
  3. Check which resolution setting a Gemini score used. Gemini 3 Pro scores 72.7% at extra_high and 60.5% at high. When DeepMind published the PDF, extra_high was not yet available in the API, so the number you could reproduce was the lower one.
  4. Budget calls per screenshot, not calls per task. The top leaderboard entry, GroundCascade at 85.1%, makes about 2.6 model calls per screenshot. Zoom, cascade and MVP wrappers lift open 8B to 32B models by 5.6 to 9.6 points over the same weights. For a high-volume agent, an open grounding model with a wrapper is a serious option against a frontier API.
  5. Small models with a refinement loop match large ones without it. Holo2-8B with three-step agentic localization scores 71.4%, level with Holo2-235B single-step at 70.6%. If your latency budget allows three passes, the 8B model reaches the 235B model's single-step accuracy.
  6. Send screenshots at native resolution where the model allows it. Opus 4.7 accepts images up to 2,576 pixels on the long edge, and ScreenSpot-Pro screens go up to 5120x2880. Check what your pipeline does to 4K and 5K captures before blaming the model.
  7. Ignore claims without numbers. OpenAI says GPT-6 Astra is state-of-the-art on ScreenSpot Pro and has published nothing to check. Run your own screenshots through each candidate with the same crop tool and the same test-time compute budget.

The public numbers mostly measure the wrapper, so the test that counts is your own. Capture screens from the applications your agent drives, label the targets, and score each candidate under the harness you will deploy. Klu runs that kind of evaluation on your own workflow, as described for operations teams, and the LLM leaderboard tracks public results as new frontier models ship.