ScreenSpot-Pro — No-tools condition of ScreenSpot-Pro GUI-grounding benchmark; distinct from with-tools condition.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.8 | Anthropic | 87.9 | 2 | 2026-08-23 |
| 2 | GPT-5.2 | OpenAI | 86.3 | 1 | 2026-08-23 |
| 3 | Qwen3.8 Max Preview | Alibaba | 84.5 | 1 | 2026-08-23 |
| 4 | Muse Spark | Meta | 84.1 | 1 | 2026-08-23 |
| 5 | Claude Opus 4.7 | Anthropic | 79.5 | 3 | 2026-06-21 |
| 6 | Qwen3.7 Plus Preview | Alibaba | 79 | 1 | 2026-08-23 |
| 7 | muse-glimmer-30b | Meta | 75.4 | 1 | 2026-08-23 |
| 8 | Gemini 3 Pro | 72.7 | 1 | 2026-08-23 | |
| 9 | Qwen3.5 122B A10B | Alibaba | 70.4 | 1 | 2026-08-23 |
| 10 | Qwen3.5 27B | Alibaba | 70.3 | 1 | 2026-08-23 |
| 11 | Gemini 3 Flash Preview | 69.1 | 1 | 2026-08-23 | |
| 12 | Qwen3.5 35B A3B | Alibaba | 68.6 | 1 | 2026-08-23 |
| 13 | Qwen3.6 Plus | Alibaba | 68.2 | 1 | 2026-08-23 |
| 14 | Qwen3 VL 235B A22B Instruct | Alibaba | 62 | 1 | 2026-08-23 |
| 15 | Qwen3 VL 235B A22B Reasoning | Alibaba | 61.8 | 1 | 2026-08-23 |
| 16 | Qwen3 VL 30B A3B Instruct | Alibaba | 60.5 | 1 | 2026-08-23 |
| 17 | Qwen3 VL 4B Instruct | Alibaba | 59.5 | 1 | 2026-08-23 |
| 18 | Qwen3 VL 32B Instruct | Alibaba | 57.9 | 1 | 2026-08-23 |
| 19 | Claude Opus 4.6 | Anthropic | 57.7 | 2 | 2026-06-21 |
| 20 | Qwen3 VL 30B A3B Reasoning | Alibaba | 57.3 | 1 | 2026-08-23 |
| 21 | Qwen3 VL 32B Reasoning | Alibaba | 57.1 | 1 | 2026-08-23 |
| 22 | Qwen3 VL 8B Instruct | Alibaba | 54.6 | 1 | 2026-08-23 |
| 23 | Step3 VL 10B | StepFun | 51.5 | 2 | 2026-06-05 |
| 24 | Qwen3 VL 4B (Reasoning) | Alibaba | 49.2 | 1 | 2026-08-23 |
| 25 | Qwen3 VL Thinking (8B) | Alibaba | 46.6 | 3 | 2026-08-23 |
| 26 | GLM-4.6V-Flash (9B) | Z.ai | 45.7 | 2 | 2026-06-05 |
| 27 | Qwen2.5 VL 72B | Alibaba | 43.6 | 1 | 2026-08-23 |
| 28 | Qwen2.5 VL 32B | Alibaba | 39.4 | 1 | 2026-08-23 |
| 29 | MiMo VL RL 2508 (7B) | Xiaomi | 34.8 | 2 | 2026-06-05 |
| 30 | InternVL-3.5 (8B) | OpenGVLab | 15.4 | 2 | 2026-06-05 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.