MMBench, English subset — StepFun source; not merged w/ MMBench-EN-v1.1 given no shared source and unspecified version here.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Gemini 2.5 Pro | 93.2 | 2 | 2026-06-05 | |
| 2 | GLM-4.6V (106B-A12B) | Z.ai | 92.8 | 2 | 2026-06-05 |
| 3 | Qwen3 VL 235B A22B Reasoning | Alibaba | 92.7 | 1 | 2026-06-05 |
| 4 | Step3 VL 10B | StepFun | 92.4 | 2 | 2026-06-05 |
| 5 | GLM-4.6V-Flash (9B) | Z.ai | 91 | 2 | 2026-06-05 |
| 6 | Qwen3 VL Thinking (8B) | Alibaba | 90.5 | 2 | 2026-06-05 |
| 7 | MiMo VL RL 2508 (7B) | Xiaomi | 89.9 | 2 | 2026-06-05 |
| 8 | InternVL-3.5 (8B) | OpenGVLab | 88.2 | 2 | 2026-06-05 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.