LongVideoBench — Diff source/roster/range (65.6-79.8) than LongVideoBench (val) (51.5-66.7) -- kept separate.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Kimi K2.5 | Moonshot | 79.8 | 2 | 2026-08-23 |
| 2 | Gemini 3 Pro | 77.7 | 1 | 2026-06-15 | |
| 3 | GPT-5.2 | OpenAI | 76.5 | 1 | 2026-06-15 |
| 4 | Claude Opus 4.5 | Anthropic | 67.2 | 1 | 2026-08-24 |
| 5 | Qwen3 VL 235B A22B Reasoning | Alibaba | 65.6 | 1 | 2026-06-15 |
| 6 | Mage-VL-4B | Microsoft | 61.3 | 1 | 2026-07-26 |
| 7 | Qwen3 VL 4B (Reasoning) | Alibaba | 57.7 | 1 | 2026-07-26 |
| 8 | Phi 4 R V 15B | Microsoft | 51.2 | 1 | 2026-07-26 |
| 9 | Phi 4 MM 5.6B | Microsoft | 41.1 | 1 | 2026-07-26 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.