| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Gemini 1.5 Pro | 62.6 | 1 | 2026-08-24 | |
| 2 | GPT-4o mini | OpenAI | 61.2 | 1 | 2026-08-24 |
| 3 | Claude 3.5 Sonnet | Anthropic | 55.9 | 1 | 2026-08-24 |
| 4 | InternVL-2-8B | OpenGVLab | 52.6 | 1 | 2026-08-24 |
| 5 | Phi 3.5 Vision Instruct | Microsoft | 50.8 | 1 | 2026-08-24 |
| 6 | LlaVA-Interleave-Qwen-7B | Alibaba | 50.2 | 1 | 2026-08-24 |
| 7 | InternVL-2-4B | OpenGVLab | 49.9 | 1 | 2026-08-24 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.