ChartQA — Standard chart-understanding VQA benchmark; base variant
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | MiniMax VL 01 | MiniMax | 91.7 | 1 | 2026-06-05 |
| 2 | InternVL2.5-78B | OpenGVLab | 91.5 | 1 | 2026-06-05 |
| 3 | Claude 3.5 Sonnet | Anthropic | 90.8 | 2 | 2026-08-23 |
| 4 | Llama 4 Maverick | Meta | 90 | 2 | 2026-08-23 |
| 5 | Qwen2.5 VL 72B | Alibaba | 89.5 | 1 | 2026-08-23 |
| 6 | Nova Pro (Non-Reasoning) | Amazon | 89.2 | 1 | 2026-08-23 |
| 7 | Llama 4 Scout | Meta | 88.8 | 3 | 2026-08-23 |
| 8 | Gemini 1.5 Pro | 88.7 | 1 | 2026-06-05 | |
| 9 | Gemini 2.0 Flash Exp | 88.3 | 1 | 2026-06-05 | |
| 10 | Qwen2-VL-72B-Instruct | Alibaba | 88.3 | 1 | 2026-08-24 |
| 11 | GPT-4o | OpenAI | 88.1 | 2 | 2026-08-23 |
| 12 | Pixtral Large | Mistral | 88.1 | 1 | 2026-08-23 |
| 13 | Nova Lite (Non-Reasoning) | Amazon | 86.8 | 1 | 2026-08-23 |
| 14 | Llama 3.2 Instruct 90B (Vision) | Meta | 85.5 | 1 | 2026-06-05 |
| 15 | Qwen2.5 Omni 7B | Alibaba | 85.3 | 1 | 2026-08-23 |
| 16 | Mage-VL-4B | Microsoft | 84.9 | 1 | 2026-07-26 |
| 17 | Qwen3 VL 4B (Reasoning) | Alibaba | 84 | 1 | 2026-07-26 |
| 18 | Phi 4 MM 5.6B | Microsoft | 83.8 | 1 | 2026-07-26 |
| 19 | Phi 4 R V 15B | Microsoft | 83.4 | 1 | 2026-07-26 |
| 20 | Phi 3.5 Vision Instruct | Microsoft | 81.8 | 1 | 2026-08-23 |
| 21 | Phi 4 Multimodal Instruct | Microsoft | 81.4 | 1 | 2026-08-23 |
| 22 | LFM2.5-VL-3B | Liquid AI | 81.3 | 2 | 2026-08-23 |
| 23 | North-Micro-Vision-Instruct | Cohere | 80.8 | 1 | 2026-08-23 |
| 24 | LFM2-VL-3B (3.1B) | Liquid AI | 80.4 | 1 | 2026-08-12 |
| 25 | Gemma 3 27B | 78 | 1 | 2026-08-23 | |
| 26 | Gemma 3 12B | 75.7 | 1 | 2026-08-23 | |
| 27 | Gemma 3 4B | 68.8 | 1 | 2026-08-23 | |
| 28 | Gemma 4 E2B | 43.5 | 1 | 2026-08-12 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.