VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Claude 3.5 Sonnet or InternVL-2-4B?
Across 10 shared benchmarks, Claude 3.5 Sonnet scores higher on 6 and InternVL-2-4B on 3, with 1 level. The widest gap is MMMU (val) (Pass@1), where Claude 3.5 Sonnet scores 68.3 against 44.2.

Claude 3.5 Sonnet vs InternVL-2-4B

Across 10 shared benchmarks, Claude 3.5 Sonnet scores higher on 6 and InternVL-2-4B on 3, with 1 level. The widest gap is MMMU (val) (Pass@1), where Claude 3.5 Sonnet scores 68.3 against 44.2.

AnthropicvsOpenGVLab10 shared benchmarks63 head-to-head
BenchmarkClaude 3.5 SonnetInternVL-2-4B
AI2D94.777.3
BLINK56.545.9
InterGPS (test)45.645.6
MathVista67.753.7
MMBench (dev-en)82.383.4
MMMU (val) (Pass@1)68.344.2
POPE (test)76.683.3
ScienceQA (img-test)73.894.9
TextVQA (val)70.566.2
Video-MME Overall55.949.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.