MultiPL-E — Explicit MBPP-derived subset, distinct from HumanEval-derived and unscoped.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Llama 3.1 Instruct 405B | Meta | 65.7 | 1 | 2026-08-23 |
| 2 | Llama 3.1 70B Instruct | Meta | 62 | 1 | 2026-08-23 |
| 3 | Kimi K2 Base | Moonshot | 58.8 | 4 | 2026-06-13 |
| 4 | Step3.5 Flash Base | StepFun | 58 | 2 | 2026-06-12 |
| 5 | MiMo V2 Flash Base | Xiaomi | 56.7 | 4 | 2026-06-13 |
| 6 | DeepSeek-V3.1-Base | DeepSeek | 52.5 | 4 | 2026-06-13 |
| 7 | Llama 3.1 8B Instruct | Meta | 52.4 | 1 | 2026-08-23 |
| 8 | DeepSeek-V3.2-Exp-Base | DeepSeek | 50.6 | 4 | 2026-06-13 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.