Across 12 models scored on InfoVQA_test, Kimi K2.5 leads at 92.6, ahead of Qwen3 VL 235B A22B Reasoning at 89.5. The median tracked score is 86, and the field spans 80.3 to 92.6.
Data as of August 25, 2026| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Kimi K2.5 | Moonshot | 92.6 | 1 | 2026-08-25 |
| 2 | Qwen3 VL 235B A22B Reasoning | Alibaba | 89.5 | 1 | 2026-08-25 |
| 3 | Qwen3 VL 32B Reasoning | Alibaba | 89.2 | 1 | 2026-08-25 |
| 4 | Qwen3 VL 235B A22B Instruct | Alibaba | 89.2 | 1 | 2026-08-25 |
| 5 | Qwen3 VL 32B Instruct | Alibaba | 87 | 1 | 2026-08-25 |
| 6 | Qwen3 VL 30B A3B Reasoning | Alibaba | 86 | 1 | 2026-08-25 |
| 7 | Qwen3 VL Thinking (8B) | Alibaba | 86 | 1 | 2026-08-25 |
| 8 | Qwen2 VL 72B Instruct | Alibaba | 84.5 | 2 | 2026-08-25 |
| 9 | Qwen3 VL 8B Instruct | Alibaba | 83.1 | 1 | 2026-08-25 |
| 10 | Qwen3 VL 4B (Reasoning) | Alibaba | 83 | 1 | 2026-08-25 |
| 11 | Qwen3 VL 30B A3B Instruct | Alibaba | 82 | 1 | 2026-08-25 |
| 12 | Qwen3 VL 4B Instruct | Alibaba | 80.3 | 1 | 2026-08-25 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.