Across 3 shared benchmarks, GPT-4o scores higher on 3 and InternVL-2-4B on 0. The widest gap is MMMU (val) (Pass@1), where GPT-4o scores 69.1 against 44.2.
| Benchmark | GPT-4o | InternVL-2-4B |
|---|---|---|
| AI2D | 94.2 | 77.3 |
| MathVista | 63.8 | 53.7 |
| MMMU (val) (Pass@1) | 69.1 | 44.2 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.