Across 13 shared benchmarks, DeepSeek-V3 scores higher on 13 and Mistral 7B on 0. The widest gap is MATH, where DeepSeek-V3 scores 91.2 against 13.1.
| Benchmark | DeepSeek-V3 | Mistral 7B |
|---|---|---|
| arc_challenge | 95.3 | 60 |
| bbh | 87.5 | 56.1 |
| C-Eval | 90.1 | 47.4 |
| GPQA Diamond | 68.4 | 24.7 |
| GSM8K | 96.7 | 52.2 |
| hellaswag | 88.9 | 83.2 |
| humaneval | 92.1 | 29.3 |
| MATH | 91.2 | 13.1 |
| MBPP | 75.4 | 51.1 |
| mmlu | 89.4 | 64.2 |
| MMLU-Pro | 81.2 | 30.9 |
| MultiPL-E | 83.1 | 29.4 |
| winogrande | 84.9 | 78.4 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.