Across 16 shared benchmarks, Claude 3.5 Sonnet scores higher on 4 and Qwen2 VL 72B Instruct on 12. The widest gap is VCR_en easy, where Qwen2 VL 72B Instruct scores 91.9 against 63.9.
| Benchmark | Claude 3.5 Sonnet | Qwen2 VL 72B Instruct |
|---|---|---|
| chartqa | 90.8 | 88.3 |
| docvqa | 95.2 | 96.5 |
| DocVQA_test | 95.2 | 96.5 |
| HallBench_avg | 49.9 | 58.1 |
| MathVista | 67.7 | 70.5 |
| MMBench-CN_test | 80.7 | 86.6 |
| MMBench-EN_test | 79.7 | 86.5 |
| MMBench-V1.1_test | 78.5 | 85.9 |
| MME_sum | 1920 | 2483 |
| MMMU | 72 | 64.5 |
| MMMU (val) (Pass@1) | 68.3 | 64.5 |
| MMMU-Pro | 54.7 | 46.2 |
| MMStar | 62.2 | 68.3 |
| OCRBench | 790 | 877 |
| RealWorldQA | 60.1 | 77.8 |
| VCR_en easy | 63.9 | 91.9 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.