MM MT-Bench (multimodal multi-turn benchmark) — Distinct multimodal variant of MT-Bench, unrelated to MMLU/MMAU/MMAR/MMBench family.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Mistral Large 3 | Mistral | 84.9 | 1 | 2026-08-23 |
| 2 | Qwen3 VL 235B A22B Instruct | Alibaba | 8.5 | 1 | 2026-08-23 |
| 3 | Qwen3 VL 235B A22B Reasoning | Alibaba | 8.5 | 1 | 2026-08-23 |
| 4 | Ministral 3 14B | Mistral | 8.5 | 4 | 2026-06-12 |
| 5 | Qwen3 VL 32B Instruct | Alibaba | 8.4 | 1 | 2026-08-23 |
| 6 | Qwen3 VL 32B Reasoning | Alibaba | 8.3 | 1 | 2026-08-23 |
| 7 | Qwen3 VL 30B A3B Instruct | Alibaba | 8.1 | 1 | 2026-08-23 |
| 8 | Ministral 3 8B | Mistral | 8.1 | 4 | 2026-06-12 |
| 9 | Qwen3 VL 4B Instruct | Alibaba | 8 | 5 | 2026-08-23 |
| 10 | Qwen3 VL Thinking (8B) | Alibaba | 8 | 1 | 2026-08-23 |
| 11 | Qwen3 VL 8B Instruct | Alibaba | 8 | 5 | 2026-08-23 |
| 12 | Qwen3 VL 30B A3B Reasoning | Alibaba | 7.9 | 1 | 2026-08-23 |
| 13 | Ministral 3 3B | Mistral | 7.8 | 4 | 2026-06-12 |
| 14 | Qwen3 VL 4B (Reasoning) | Alibaba | 7.7 | 1 | 2026-08-23 |
| 15 | Gemma 3 12B | 6.7 | 4 | 2026-06-05 | |
| 16 | Gemma 3 4B | 5.2 | 4 | 2026-06-05 | |
| 17 | Pixtral Large | Mistral | 0.7 | 1 | 2026-08-23 |
| 18 | Qwen2.5 Omni 7B | Alibaba | 0.1 | 1 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.