Across 10 shared benchmarks, Claude 3.5 Sonnet scores higher on 6 and InternVL-2-8B on 4. The widest gap is MMMU (val) (Pass@1), where Claude 3.5 Sonnet scores 68.3 against 46.3.
| Benchmark | Claude 3.5 Sonnet | InternVL-2-8B |
|---|---|---|
| AI2D | 94.7 | 81.4 |
| BLINK | 56.5 | 45.4 |
| InterGPS (test) | 45.6 | 53.2 |
| MathVista | 67.7 | 51.1 |
| MMBench (dev-en) | 82.3 | 87 |
| MMMU (val) (Pass@1) | 68.3 | 46.3 |
| POPE (test) | 76.6 | 84.2 |
| ScienceQA (img-test) | 73.8 | 95.9 |
| TextVQA (val) | 70.5 | 68.8 |
| Video-MME Overall | 55.9 | 52.6 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.