MM-Vet — Same MM-Vet, close band + overlapping models, but declared metric differs; kept apart.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen2.5 VL 72B | Alibaba | 76.2 | 1 | 2026-08-23 |
| 2 | GPT-4o | OpenAI | 69.1 | 1 | 2026-06-05 |
| 3 | GPT-4o mini | OpenAI | 66.9 | 1 | 2026-06-05 |
| 4 | Kimi VL A3B Instruct | Moonshot | 66.7 | 1 | 2026-06-12 |
| 5 | Gemma 3 12B | 64.9 | 1 | 2026-06-05 | |
| 6 | LFM2.5-VL-450M-Extract | Liquid AI | 41.1 | 1 | 2026-08-09 |
| 7 | LFM2-VL-450M | Liquid AI | 33.9 | 1 | 2026-08-09 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.