DocVQA — Standard document visual-question-answering benchmark
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen2.5 VL 72B | Alibaba | 96.4 | 1 | 2026-08-23 |
| 2 | MiniMax VL 01 | MiniMax | 96.4 | 1 | 2026-06-05 |
| 3 | InternVL2.5-78B | OpenGVLab | 96.1 | 1 | 2026-06-05 |
| 4 | Qwen2.5 Omni 7B | Alibaba | 95.2 | 1 | 2026-08-23 |
| 5 | Claude 3.5 Sonnet | Anthropic | 95.2 | 3 | 2026-08-23 |
| 6 | Qwen2.5 VL 32B | Alibaba | 94.8 | 1 | 2026-08-23 |
| 7 | Llama 4 Maverick | Meta | 94.4 | 2 | 2026-08-23 |
| 8 | Llama 4 Scout | Meta | 94.4 | 3 | 2026-08-23 |
| 9 | Grok 2 | SpaceXAI | 93.6 | 1 | 2026-08-23 |
| 10 | Nova Pro (Non-Reasoning) | Amazon | 93.5 | 1 | 2026-08-23 |
| 11 | Pixtral Large | Mistral | 93.3 | 1 | 2026-08-23 |
| 12 | Phi 4 Multimodal Instruct | Microsoft | 93.2 | 1 | 2026-08-23 |
| 13 | Grok 2 Mini | SpaceXAI | 93.2 | 1 | 2026-08-23 |
| 14 | Gemini 2.0 Flash Exp | 92.9 | 1 | 2026-06-05 | |
| 15 | GPT-4o | OpenAI | 92.8 | 3 | 2026-08-23 |
| 16 | Nova Lite (Non-Reasoning) | Amazon | 92.4 | 1 | 2026-08-23 |
| 17 | North-Micro-Vision-Instruct | Cohere | 92.1 | 1 | 2026-08-23 |
| 18 | Gemini 1.5 Pro | 91.5 | 2 | 2026-06-12 | |
| 19 | LFM2.5-VL-3B | Liquid AI | 91.1 | 1 | 2026-08-23 |
| 20 | Llama 3.2 Instruct 90B (Vision) | Meta | 90.1 | 1 | 2026-06-05 |
| 21 | Gemma 3 12B | 87.1 | 1 | 2026-08-23 | |
| 22 | Gemma 3 27B | 86.6 | 1 | 2026-08-23 | |
| 23 | Gemma 3 4B | 75.8 | 1 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.