VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Claude 3.5 Sonnet or Qwen2 VL 72B Instruct?
Across 16 shared benchmarks, Claude 3.5 Sonnet scores higher on 4 and Qwen2 VL 72B Instruct on 12. The widest gap is VCR_en easy, where Qwen2 VL 72B Instruct scores 91.9 against 63.9.

Claude 3.5 Sonnet vs Qwen2 VL 72B Instruct

Across 16 shared benchmarks, Claude 3.5 Sonnet scores higher on 4 and Qwen2 VL 72B Instruct on 12. The widest gap is VCR_en easy, where Qwen2 VL 72B Instruct scores 91.9 against 63.9.

AnthropicvsAlibaba16 shared benchmarks412 head-to-head
BenchmarkClaude 3.5 SonnetQwen2 VL 72B Instruct
chartqa90.888.3
docvqa95.296.5
DocVQA_test95.296.5
HallBench_avg49.958.1
MathVista67.770.5
MMBench-CN_test80.786.6
MMBench-EN_test79.786.5
MMBench-V1.1_test78.585.9
MME_sum19202483
MMMU7264.5
MMMU (val) (Pass@1)68.364.5
MMMU-Pro54.746.2
MMStar62.268.3
OCRBench790877
RealWorldQA60.177.8
VCR_en easy63.991.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.