MLVU (Multi-task Long Video Understanding) — Not merged w/ MLVU-MCQ (Acc) -- diff source/roster and MCQ-specific qualifier on the latter.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen3.7 Plus Preview | Alibaba | 87.4 | 1 | 2026-08-23 |
| 2 | Qwen3.5 122B A10B | Alibaba | 87.3 | 1 | 2026-08-23 |
| 3 | Qwen3.5 397B A17B | Alibaba | 86.7 | 1 | 2026-08-24 |
| 4 | Qwen3.6 Plus | Alibaba | 86.7 | 1 | 2026-08-23 |
| 5 | Qwen3.6 27B | Alibaba | 86.6 | 2 | 2026-08-24 |
| 6 | Qwen3.6 35B A3B | Alibaba | 86.2 | 1 | 2026-08-23 |
| 7 | Qwen3.5 27B | Alibaba | 85.9 | 2 | 2026-08-24 |
| 8 | Qwen3.5 35B A3B | Alibaba | 85.6 | 1 | 2026-08-23 |
| 9 | Qwen3 VL 235B A22B Instruct | Alibaba | 84.3 | 1 | 2026-08-23 |
| 10 | Qwen3 VL 235B A22B Reasoning | Alibaba | 83.8 | 1 | 2026-08-23 |
| 11 | Claude Opus 4.5 | Anthropic | 81.7 | 1 | 2026-08-24 |
| 12 | Gemini 3 Pro | 80.7 | 1 | 2026-06-15 | |
| 13 | Qwen3.5 9B | Alibaba | 52.4 | 1 | 2026-06-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.