VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

InternVL-3-78B vs Qwen2.5 VL 72B

26 SHARED BENCHMARKS

Across 26 shared benchmarks, InternVL-3-78B scores higher on 6 and Qwen2.5 VL 72B on 20. The widest gap is OCRVQA_TEST, where Qwen2.5 VL 72B scores 66.8 against 35.6.

OPENGVLABVSALIBABA26 SHARED620 HEAD-TO-HEAD
A-Bench_VAL75.979.2
AI2D83.588.4
BLINK51.963
CCBench70.873.7
ChartQAPro44.445.3
DocVQA-val83.895.8
InHouse Dataset A41.557.2
InHouse Dataset B42.659.7
Mathverse49.355.2
mathvision34.839.3
MathVista70.174.8
MMStar66.170.8
MMT-Bench_VAL63.769.5
MTVQA_TEST27.631.5
OCRVQA_TEST35.666.8
POPE88.983.3
Q-Bench1_VAL7879.9
RealWorldQA74.373.9
ScienceQA_TEST97.292.5
ScienceQA_VAL95.191.3
SEEDBench_IMG77.578.3
SEEDBench2_Plus68.573.3
TextVQA-val83.583.3

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.