VSI-Bench — Distinct named visual-spatial-intelligence benchmark; raws "VSI-Bench"/"VSIBench" both present
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Mage-VL-4B | Microsoft | 64.3 | 1 | 2026-07-26 |
| 2 | HY-Embodied 0.5 MoT-2B | Tencent | 60.5 | 1 | 2026-06-15 |
| 3 | Qwen3 VL 4B (Reasoning) | Alibaba | 55.2 | 2 | 2026-07-26 |
| 4 | MiMo Embodied 7B | Xiaomi | 48.5 | 1 | 2026-06-15 |
| 5 | Qwen3 VL 2B | Alibaba | 48 | 1 | 2026-06-15 |
| 6 | Kimi VL A3B Instruct | Moonshot | 37.4 | 1 | 2026-06-12 |
| 7 | GPT-4o | OpenAI | 34 | 1 | 2026-06-05 |
| 8 | Gemma 3 12B | 32.4 | 1 | 2026-06-05 | |
| 9 | Phi 4 R V 15B | Microsoft | 25.5 | 1 | 2026-07-26 |
| 10 | Phi 4 MM 5.6B | Microsoft | 24.1 | 1 | 2026-07-26 |
| 11 | Qwen3 VL Thinking (8B) | Alibaba | 6.8 | 1 | 2026-07-05 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.