MMSI-Bench (Multi-image Multimodal Spatial Intelligence) — Spatial-reasoning multi-image VLM benchmark; single-source Tencent card.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | GPT-5 (ChatGPT) | OpenAI | 40 | 1 | 2026-07-05 |
| 2 | HY-Embodied 0.5 MoT-2B | Tencent | 33.2 | 1 | 2026-06-15 |
| 3 | MiMo Embodied 7B | Xiaomi | 31.9 | 1 | 2026-06-15 |
| 4 | Qwen3 VL 4B (Reasoning) | Alibaba | 25.1 | 1 | 2026-06-15 |
| 5 | Qwen3 VL 2B | Alibaba | 23.6 | 1 | 2026-06-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.