RefSpatial-Bench — Spatial-referring benchmark for embodied/vision models.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen3.6 27B | Alibaba | 70 | 2 | 2026-08-24 |
| 2 | Qwen3.5 27B | Alibaba | 67.7 | 2 | 2026-08-24 |
| 3 | MiMo Embodied 7B | Xiaomi | 48 | 1 | 2026-06-15 |
| 4 | HY-Embodied 0.5 MoT-2B | Tencent | 45.8 | 1 | 2026-06-15 |
| 5 | Qwen3 VL 4B (Reasoning) | Alibaba | 45.3 | 1 | 2026-06-15 |
| 6 | Qwen3 VL 2B | Alibaba | 28.9 | 1 | 2026-06-15 |
| 7 | Gemma 4 31B | 4.7 | 1 | 2026-08-24 | |
| 8 | Qwen3 VL 235B A22B Reasoning | Alibaba | 0.7 | 1 | 2026-08-23 |
| 9 | Qwen3.5 122B A10B | Alibaba | 0.7 | 1 | 2026-08-23 |
| 10 | Qwen3.6 35B A3B | Alibaba | 0.6 | 1 | 2026-08-23 |
| 11 | Qwen3.5 35B A3B | Alibaba | 0.6 | 1 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.