Across 9 shared benchmarks, DeepSeek-V3 scores higher on 9 and Gemma 3 1B on 0. The widest gap is MMLU-Pro, where DeepSeek-V3 scores 81.2 against 14.7.
| Benchmark | DeepSeek-V3 | Gemma 3 1B |
|---|---|---|
| bbh | 87.5 | 39.1 |
| GPQA Diamond | 68.4 | 19.2 |
| GSM8K | 96.7 | 62.8 |
| humaneval | 92.1 | 41.5 |
| ifeval | 87.3 | 80.2 |
| LiveCodeBench | 49.6 | 1.9 |
| MBPP | 75.4 | 0.4 |
| MMLU-Pro | 81.2 | 14.7 |
| simpleqa | 27.7 | 2.2 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.