Across 13 models scored on DocVQA_test, Qwen3 VL 235B A22B Instruct leads at 97.1, ahead of Qwen3 VL 32B Instruct at 96.9. The median tracked score is 95.3, and the field spans 92.8 to 97.1.
Data as of August 25, 2026| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen3 VL 235B A22B Instruct | Alibaba | 97.1 | 1 | 2026-08-25 |
| 2 | Qwen3 VL 32B Instruct | Alibaba | 96.9 | 1 | 2026-08-25 |
| 3 | Qwen2 VL 72B Instruct | Alibaba | 96.5 | 2 | 2026-08-25 |
| 4 | Qwen3 VL 235B A22B Reasoning | Alibaba | 96.5 | 1 | 2026-08-25 |
| 5 | Qwen3 VL 8B Instruct | Alibaba | 96.1 | 1 | 2026-08-25 |
| 6 | Qwen3 VL 32B Reasoning | Alibaba | 96.1 | 1 | 2026-08-25 |
| 7 | Qwen3 VL Thinking (8B) | Alibaba | 95.3 | 1 | 2026-08-25 |
| 8 | Qwen3 VL 4B Instruct | Alibaba | 95.3 | 1 | 2026-08-25 |
| 9 | Claude 3.5 Sonnet | Anthropic | 95.2 | 1 | 2026-08-24 |
| 10 | Qwen3 VL 30B A3B Instruct | Alibaba | 95 | 1 | 2026-08-25 |
| 11 | Qwen3 VL 30B A3B Reasoning | Alibaba | 95 | 1 | 2026-08-25 |
| 12 | Qwen3 VL 4B (Reasoning) | Alibaba | 94.2 | 1 | 2026-08-25 |
| 13 | GPT-4o | OpenAI | 92.8 | 1 | 2026-08-24 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.