Across 15 shared benchmarks, GPT-4o scores higher on 12 and Qwen2.5 Omni 7B on 3. The widest gap is GPQA Diamond, where GPT-4o scores 70.1 against 30.8. Qwen2.5 Omni 7B is the cheaper of the two on tracked API pricing ($0.10 against $5.00 per million input tokens).
| Benchmark | GPT-4o | Qwen2.5 Omni 7B |
|---|---|---|
| AI2D | 94.2 | 83.2 |
| chartqa | 88.1 | 85.3 |
| docvqa | 92.8 | 95.2 |
| EgoSchema | 72.2 | 68.6 |
| GPQA Diamond | 70.1 | 30.8 |
| GSM8K | 95.6 | 88.7 |
| humaneval | 90.6 | 78.7 |
| mathvision | 30.4 | 25 |
| MathVista | 63.8 | 67.9 |
| mmlu_redux | 88 | 71 |
| MMLU-Pro | 74.7 | 47 |
| MMMU | 72.2 | 59.2 |
| MMMU-Pro | 59.9 | 36.6 |
| MMStar | 63.9 | 64 |
| RealWorldQA | 75.4 | 70.3 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, alibaba-official), otherwise the lowest tracked offer.