Across 9 shared benchmarks, Gemma 3 1B scores higher on 0 and Qwen2.5 Instruct 72B on 9. The widest gap is MMLU-Pro, where Qwen2.5 Instruct 72B scores 71.6 against 14.7.
| Benchmark | Gemma 3 1B | Qwen2.5 Instruct 72B |
|---|---|---|
| bbh | 39.1 | 79.8 |
| GPQA Diamond | 19.2 | 49.1 |
| GSM8K | 62.8 | 95.8 |
| humaneval | 41.5 | 86.6 |
| ifeval | 80.2 | 87.2 |
| LiveCodeBench | 1.9 | 55.5 |
| MBPP | 0.4 | 72.6 |
| MMLU-Pro | 14.7 | 71.6 |
| simpleqa | 2.2 | 10.3 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.