CV-Bench — Vision-centric spatial-reasoning benchmark for VLMs
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | HY-Embodied 0.5 MoT-2B | Tencent | 89.2 | 1 | 2026-06-15 |
| 2 | MiMo Embodied 7B | Xiaomi | 88.8 | 1 | 2026-06-15 |
| 3 | Mage-VL-4B | Microsoft | 87.8 | 1 | 2026-07-26 |
| 4 | Qwen3 VL 4B (Reasoning) | Alibaba | 85.7 | 2 | 2026-07-26 |
| 5 | Phi 4 R V 15B | Microsoft | 81.3 | 1 | 2026-07-26 |
| 6 | Qwen3 VL 2B | Alibaba | 80 | 1 | 2026-06-15 |
| 7 | Phi 4 MM 5.6B | Microsoft | 57.1 | 1 | 2026-07-26 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.