ScreenSpot Pro

GUI Grounding for Professional High-Resolution Computer Use

Models scored

23

evaluated

Modality

multimodal

Category

grounding

+3 more

Published

2025

arxiv.org

Citations

211

Semantic Scholar

Influential

58

citations

References

29

cited works

Venue

ACM Multimedia

published in

Abstract

Kaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo, et al. (+4)

Recent advancements in Multi-modal Large Language Models (MLLMs) have led to significant progress in developing GUI agents for general tasks such as web browsing and mobile phone use. However, their application in professional domains remains under-explored. These specialized workflows introduce unique challenges for GUI perception models, including high-resolution displays and complex environments which lead to smaller target sizes. In this paper, we introduce ScreenSpot-Pro, a new benchmark designed to rigorously evaluate the grounding capabilities of MLLMs in high-resolution professional settings. The benchmark comprises authentic high-resolution images from a variety of professional domains with expert annotations. It spans 23 applications across five industries and three operating systems. Existing GUI grounding models perform poorly on this dataset, with the best model achieving only 18.9%. Our experiments reveal that strategically reducing the search area enhances accuracy. Based on this insight, we propose ScreenSeekeR, a visual search method that utilizes the GUI knowledge of a strong planner to guide a cascaded search, achieving state-of-the-art performance with 48.1% without any additional training. We hope that our benchmark and findings will advance the development of GUI agents for professional settings.

groundingmultimodalspatial reasoningvision

Search

#ModelLabScore
01Claude Opus 4.8Anthropic88
02GPT-5.2OpenAI86
03Muse SparkMeta84
04Qwen3.7-PlusAlibaba Cloud / Qwen Team79
05Gemini 3 ProGoogle73
06Qwen3.5-122B-A10BAlibaba Cloud / Qwen Team70
07Qwen3.5-27BAlibaba Cloud / Qwen Team70
08Gemini 3 FlashGoogle69
09Qwen3.5-35B-A3BAlibaba Cloud / Qwen Team69
10Qwen3.6 PlusAlibaba Cloud / Qwen Team68
11Qwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team62
12Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team62
13Qwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team61
14Qwen3 VL 4B InstructAlibaba Cloud / Qwen Team60
15Qwen3 VL 32B InstructAlibaba Cloud / Qwen Team58
16Qwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team57
17Qwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team57
18Qwen3 VL 8B InstructAlibaba Cloud / Qwen Team55
19Qwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team49
20Qwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team47
21Qwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team44
22Qwen2.5 VL 32B InstructAlibaba Cloud / Qwen Team39
23Qwen2.5 VL 7B InstructAlibaba Cloud / Qwen Team29

23 of 23 models · score normalized 0–100 where available

© 2026 NYSGPT2525 LLC