Across 3 shared benchmarks, GPT-4o scores higher on 3 and InternVL-2-8B on 0. The widest gap is MMMU (val) (Pass@1), where GPT-4o scores 69.1 against 46.3.
| Benchmark | GPT-4o | InternVL-2-8B |
|---|---|---|
| AI2D | 94.2 | 81.4 |
| MathVista | 63.8 | 51.1 |
| MMMU (val) (Pass@1) | 69.1 | 46.3 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.