| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen2-VL-72B-Instruct | Alibaba | 77.9 | 1 | 2026-08-24 |
| 2 | Qwen2.5 VL 72B | Alibaba | 76.2 | 1 | 2026-08-23 |
| 3 | Grok 3 Beta | SpaceXAI | 74.5 | 1 | 2026-07-13 |
| 4 | Grok 3 Mini Beta | SpaceXAI | 74.3 | 1 | 2026-07-13 |
| 5 | GPT-4o | OpenAI | 72.2 | 1 | 2026-08-23 |
| 6 | Nova Pro (Non-Reasoning) | Amazon | 72.1 | 1 | 2026-08-23 |
| 7 | Gemini 2.0 Flash (Reasoning) | 71.5 | 1 | 2026-08-23 | |
| 8 | Nova Lite (Non-Reasoning) | Amazon | 71.4 | 1 | 2026-08-23 |
| 9 | Qwen2.5 Omni 7B | Alibaba | 68.6 | 1 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.