Models scored
23
evaluated
Modality
multimodal
Category
grounding
+3 more
Published
2025
arxiv.org
Citations
211
Semantic Scholar
Influential
58
citations
References
29
cited works
Venue
ACM Multimedia
published in
Abstract
Kaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo, et al. (+4)
Recent advancements in Multi-modal Large Language Models (MLLMs) have led to significant progress in developing GUI agents for general tasks such as web browsing and mobile phone use. However, their application in professional domains remains under-explored. These specialized workflows introduce unique challenges for GUI perception models, including high-resolution displays and complex environments which lead to smaller target sizes. In this paper, we introduce ScreenSpot-Pro, a new benchmark designed to rigorously evaluate the grounding capabilities of MLLMs in high-resolution professional settings. The benchmark comprises authentic high-resolution images from a variety of professional domains with expert annotations. It spans 23 applications across five industries and three operating systems. Existing GUI grounding models perform poorly on this dataset, with the best model achieving only 18.9%. Our experiments reveal that strategically reducing the search area enhances accuracy. Based on this insight, we propose ScreenSeekeR, a visual search method that utilizes the GUI knowledge of a strong planner to guide a cascaded search, achieving state-of-the-art performance with 48.1% without any additional training. We hope that our benchmark and findings will advance the development of GUI agents for professional settings.
Search
| # | Model | Lab | Score |
|---|---|---|---|
| 01 | Claude Opus 4.8 | Anthropic | 88 |
| 02 | GPT-5.2 | OpenAI | 86 |
| 03 | Muse Spark | Meta | 84 |
| 04 | Qwen3.7-Plus | Alibaba Cloud / Qwen Team | 79 |
| 05 | Gemini 3 Pro | 73 | |
| 06 | Qwen3.5-122B-A10B | Alibaba Cloud / Qwen Team | 70 |
| 07 | Qwen3.5-27B | Alibaba Cloud / Qwen Team | 70 |
| 08 | Gemini 3 Flash | 69 | |
| 09 | Qwen3.5-35B-A3B | Alibaba Cloud / Qwen Team | 69 |
| 10 | Qwen3.6 Plus | Alibaba Cloud / Qwen Team | 68 |
| 11 | Qwen3 VL 235B A22B Instruct | Alibaba Cloud / Qwen Team | 62 |
| 12 | Qwen3 VL 235B A22B Thinking | Alibaba Cloud / Qwen Team | 62 |
| 13 | Qwen3 VL 30B A3B Instruct | Alibaba Cloud / Qwen Team | 61 |
| 14 | Qwen3 VL 4B Instruct | Alibaba Cloud / Qwen Team | 60 |
| 15 | Qwen3 VL 32B Instruct | Alibaba Cloud / Qwen Team | 58 |
| 16 | Qwen3 VL 30B A3B Thinking | Alibaba Cloud / Qwen Team | 57 |
| 17 | Qwen3 VL 32B Thinking | Alibaba Cloud / Qwen Team | 57 |
| 18 | Qwen3 VL 8B Instruct | Alibaba Cloud / Qwen Team | 55 |
| 19 | Qwen3 VL 4B Thinking | Alibaba Cloud / Qwen Team | 49 |
| 20 | Qwen3 VL 8B Thinking | Alibaba Cloud / Qwen Team | 47 |
| 21 | Qwen2.5 VL 72B Instruct | Alibaba Cloud / Qwen Team | 44 |
| 22 | Qwen2.5 VL 32B Instruct | Alibaba Cloud / Qwen Team | 39 |
| 23 | Qwen2.5 VL 7B Instruct | Alibaba Cloud / Qwen Team | 29 |
23 of 23 models · score normalized 0–100 where available