Across 14 shared benchmarks, Mistral 7B scores higher on 1 and Qwen2 Instruct (72B) on 13. The widest gap is MATH, where Qwen2 Instruct (72B) scores 79 against 13.1.
| Benchmark | Mistral 7B | Qwen2 Instruct (72B) |
|---|---|---|
| arc_challenge | 60 | 68.9 |
| bbh | 56.1 | 82.4 |
| C-Eval | 47.4 | 83.8 |
| GPQA Diamond | 24.7 | 42.4 |
| GSM8K | 52.2 | 92 |
| hellaswag | 83.2 | 87.6 |
| humaneval | 29.3 | 86 |
| MATH | 13.1 | 79 |
| MBPP | 51.1 | 0.8 |
| mmlu | 64.2 | 76.9 |
| MMLU-Pro | 30.9 | 64.4 |
| MultiPL-E | 29.4 | 69.2 |
| TruthfulQA | 42.2 | 54.8 |
| winogrande | 78.4 | 85.1 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.