OSWorld-Verified — the OSWorld team's 2025 corrected revision of the computer-use suite; the current standard track (scores not comparable to original OSWorld). Carries the scored axis.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen3.8 Max Preview | Alibaba | 86.1 | 1 | 2026-08-23 |
| 2 | Claude Mythos 5 | Anthropic | 85.4 | 3 | 2026-08-23 |
| 3 | Claude Fable 5 | Anthropic | 85 | 4 | 2026-08-23 |
| 4 | Kimi K3 | Moonshot | 84.8 | 1 | 2026-07-27 |
| 5 | Qwen3.8 27B | Alibaba | 84.3 | 2 | 2026-08-24 |
| 6 | Claude Opus 4.8 | Anthropic | 83.4 | 6 | 2026-08-24 |
| 7 | GPT-5.6 Sol | OpenAI | 83 | 1 | 2026-08-24 |
| 8 | Gemini 3.6 Flash | 83 | 6 | 2026-08-24 | |
| 9 | Claude Opus 4.7 | Anthropic | 82.8 | 7 | 2026-08-23 |
| 10 | Claude Sonnet 5 | Anthropic | 81.2 | 3 | 2026-08-23 |
| 11 | Muse Spark 1.1 | Meta | 80.8 | 2 | 2026-08-23 |
| 12 | GPT-5.5 | OpenAI | 78.7 | 5 | 2026-08-23 |
| 13 | Claudes | Anthropic | 78.5 | 1 | 2026-08-12 |
| 14 | Claude Sonnet 4.6 | Anthropic | 78.5 | 3 | 2026-08-23 |
| 15 | Gemini 3.5 Flash Cyber | 78.4 | 1 | 2026-08-04 | |
| 16 | Gemini 3.5 Flash | 78.4 | 8 | 2026-08-24 | |
| 17 | Gemini 3.1 Pro | 76.2 | 4 | 2026-07-06 | |
| 18 | GPT-5.4 | OpenAI | 75 | 4 | 2026-08-23 |
| 19 | Gemini 3.5 Flash Lite | 74 | 5 | 2026-08-24 | |
| 20 | Qwen3.7 Plus Preview | Alibaba | 73.3 | 2 | 2026-08-24 |
| 21 | Kimi K2.6 | Moonshot | 73.1 | 2 | 2026-08-23 |
| 22 | Claude Opus 4.6 | Anthropic | 72.7 | 5 | 2026-08-24 |
| 23 | GPT-5.6 Luna | OpenAI | 72.6 | 1 | 2026-07-22 |
| 24 | GPT-5.4 mini | OpenAI | 72.1 | 1 | 2026-08-23 |
| 25 | MiniMax M3 | MiniMax | 70.1 | 1 | 2026-08-23 |
| 26 | Qwen3 VL 235B A22B Instruct | Alibaba | 66.7 | 1 | 2026-08-23 |
| 27 | Claude Opus 4.5 | Anthropic | 66.3 | 3 | 2026-08-23 |
| 28 | muse-glimmer-30b | Meta | 65.9 | 2 | 2026-08-24 |
| 29 | Gemini 3 Flash Preview | 65.1 | 4 | 2026-08-04 | |
| 30 | GPT-5.3 Codex | OpenAI | 64.7 | 1 | 2026-08-23 |
| 31 | Qwen3.6 27B | Alibaba | 63.9 | 1 | 2026-08-24 |
| 32 | Kimi K2.5 | Moonshot | 63.3 | 1 | 2026-05-03 |
| 33 | Qwen3.6 Plus | Alibaba | 62.5 | 1 | 2026-08-23 |
| 34 | GLM 5V Turbo | Z.ai | 62.3 | 1 | 2026-08-23 |
| 35 | Claude Sonnet 4.5 | Anthropic | 61.4 | 3 | 2026-08-23 |
| 36 | Qwen3.5 122B A10B | Alibaba | 58 | 1 | 2026-08-23 |
| 37 | Qwen3.5 27B | Alibaba | 56.2 | 1 | 2026-08-23 |
| 38 | Qwen3.5 35B A3B | Alibaba | 54.5 | 1 | 2026-08-23 |
| 39 | Claude Haiku 4.5 | Anthropic | 50.7 | 1 | 2026-08-23 |
| 40 | Claude Opus 4.1 | Anthropic | 44.4 | 1 | 2026-07-29 |
| 41 | Claude Sonnet 4 | Anthropic | 42.2 | 1 | 2026-08-13 |
| 42 | Qwen3 VL 32B Reasoning | Alibaba | 41 | 1 | 2026-08-23 |
| 43 | GPT-5.4 nano | OpenAI | 39 | 1 | 2026-08-23 |
| 44 | Qwen3 VL 235B A22B Reasoning | Alibaba | 38.1 | 1 | 2026-08-23 |
| 45 | Qwen3 VL 8B Instruct | Alibaba | 33.9 | 1 | 2026-08-23 |
| 46 | Qwen3 VL Thinking (8B) | Alibaba | 33.9 | 1 | 2026-08-23 |
| 47 | Qwen3 VL 32B Instruct | Alibaba | 32.6 | 1 | 2026-08-23 |
| 48 | Qwen3 VL 4B (Reasoning) | Alibaba | 31.4 | 1 | 2026-08-23 |
| 49 | Qwen3 VL 30B A3B Reasoning | Alibaba | 30.6 | 1 | 2026-08-23 |
| 50 | Qwen3 VL 30B A3B Instruct | Alibaba | 30.3 | 1 | 2026-08-23 |
| 51 | Qwen3 VL 4B Instruct | Alibaba | 26.2 | 1 | 2026-08-23 |
| 52 | Qwen2.5 VL 72B | Alibaba | 8.8 | 1 | 2026-08-23 |
| 53 | Kimi VL A3B Instruct | Moonshot | 8.2 | 1 | 2026-06-12 |
| 54 | Qwen2.5 VL 32B | Alibaba | 5.9 | 1 | 2026-08-23 |
| 55 | GPT-4o | OpenAI | 5 | 1 | 2026-06-05 |
| 56 | Opus 5 | Anthropic | 4 | 1 | 2026-07-26 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.