Across 13 shared benchmarks, Mistral 7B scores higher on 0 and Qwen2.5 Instruct 72B on 13. The widest gap is MATH, where Qwen2.5 Instruct 72B scores 88.4 against 13.1.
| Benchmark | Mistral 7B | Qwen2.5 Instruct 72B |
|---|---|---|
| arc_challenge | 60 | 94.5 |
| bbh | 56.1 | 79.8 |
| C-Eval | 47.4 | 89.2 |
| GPQA Diamond | 24.7 | 49.1 |
| GSM8K | 52.2 | 95.8 |
| hellaswag | 83.2 | 84.8 |
| humaneval | 29.3 | 86.6 |
| MATH | 13.1 | 88.4 |
| MBPP | 51.1 | 72.6 |
| mmlu | 64.2 | 86.1 |
| MMLU-Pro | 30.9 | 71.6 |
| MultiPL-E | 29.4 | 75.1 |
| winogrande | 78.4 | 82.3 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.