VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, GPT-4o or Qwen2.5 Omni 7B?
Across 15 shared benchmarks, GPT-4o scores higher on 12 and Qwen2.5 Omni 7B on 3. The widest gap is GPQA Diamond, where GPT-4o scores 70.1 against 30.8. Qwen2.5 Omni 7B is the cheaper of the two on tracked API pricing ($0.10 against $5.00 per million input tokens).

GPT-4o vs Qwen2.5 Omni 7B

Across 15 shared benchmarks, GPT-4o scores higher on 12 and Qwen2.5 Omni 7B on 3. The widest gap is GPQA Diamond, where GPT-4o scores 70.1 against 30.8. Qwen2.5 Omni 7B is the cheaper of the two on tracked API pricing ($0.10 against $5.00 per million input tokens).

OpenAIvsAlibaba15 shared benchmarks123 head-to-head
BenchmarkGPT-4oQwen2.5 Omni 7B
AI2D94.283.2
chartqa88.185.3
docvqa92.895.2
EgoSchema72.268.6
GPQA Diamond70.130.8
GSM8K95.688.7
humaneval90.678.7
mathvision30.425
MathVista63.867.9
mmlu_redux8871
MMLU-Pro74.747
MMMU72.259.2
MMMU-Pro59.936.6
MMStar63.964
RealWorldQA75.470.3

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, alibaba-official), otherwise the lowest tracked offer.