Across 20 shared benchmarks, Gemma (7B) scores higher on 14 and Mistral 7B on 6. The widest gap is MATH, where Gemma (7B) scores 50 against 13.1.
| Benchmark | Gemma (7B) | Mistral 7B |
|---|---|---|
| arc_challenge | 61.1 | 60 |
| bbh | 55.1 | 56.1 |
| C-Eval | 43.6 | 47.4 |
| EvalPlus | 39.6 | 36.4 |
| GPQA Diamond | 25.7 | 24.7 |
| GSM8K | 55.9 | 52.2 |
| hellaswag | 82.2 | 83.2 |
| humaneval | 37.2 | 29.3 |
| MATH | 50 | 13.1 |
| MBPP | 50.6 | 51.1 |
| mmlu | 64.6 | 64.2 |
| MMLU-Pro | 33.7 | 30.9 |
| Multi-Exam | 42.7 | 47.1 |
| Multi-Mathematics | 39.1 | 26.3 |
| Multi-Translation | 31.2 | 23.3 |
| Multi-Understanding | 58.3 | 63.3 |
| MultiPL-E | 29.7 | 29.4 |
| Theorem QA | 21.5 | 19.2 |
| TruthfulQA | 44.8 | 42.2 |
| winogrande | 79 | 78.4 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.