VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

InternVL-3-8B vs Qwen2.5 VL 72B

26 SHARED BENCHMARKS

Across 26 shared benchmarks, InternVL-3-8B scores higher on 5 and Qwen2.5 VL 72B on 21. The widest gap is OCRVQA_TEST, where Qwen2.5 VL 72B scores 66.8 against 39.

OPENGVLABVSALIBABA26 SHARED521 HEAD-TO-HEAD
A-Bench_VAL75.979.2
AI2D85.188.4
BLINK55.963
CCBench77.873.7
ChartQAPro37.345.3
InHouse Dataset A40.657.2
InHouse Dataset B36.359.7
Mathverse43.755.2
mathvision29.639.3
MathVista69.574.8
MMStar68.470.8
MMT-Bench_VAL65.269.5
MTVQA_TEST30.331.5
OCRVQA_TEST3966.8
POPE90.683.3
Q-Bench1_VAL7679.9
RealWorldQA71.173.9
ScienceQA_TEST9892.5
ScienceQA_VAL97.891.3
SEEDBench_IMG7778.3
SEEDBench2_Plus69.573.3
TextVQA-val82.283.3

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.