ReMI — Reasoning with Multiple Images benchmark.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Step3 VL 10B | StepFun | 67.3 | 2 | 2026-06-05 |
| 2 | MiMo VL RL 2508 (7B) | Xiaomi | 63.1 | 2 | 2026-06-05 |
| 3 | GLM-4.6V-Flash (9B) | Z.ai | 60.8 | 2 | 2026-06-05 |
| 4 | Qwen3 VL Thinking (8B) | Alibaba | 57.2 | 2 | 2026-06-05 |
| 5 | InternVL-3.5 (8B) | OpenGVLab | 52.6 | 2 | 2026-06-05 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.