MultiPL-E — Explicit HumanEval-derived subset, distinct from MBPP-derived and unscoped.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Llama 3.1 Instruct 405B | Meta | 75.2 | 1 | 2026-08-23 |
| 2 | Step3.5 Flash Base | StepFun | 67.7 | 2 | 2026-06-12 |
| 3 | Llama 3.1 70B Instruct | Meta | 65.5 | 1 | 2026-08-23 |
| 4 | Kimi K2 Base | Moonshot | 60.5 | 4 | 2026-06-13 |
| 5 | MiMo V2 Flash Base | Xiaomi | 59.5 | 4 | 2026-06-13 |
| 6 | Llama 3.1 8B Instruct | Meta | 50.8 | 1 | 2026-08-23 |
| 7 | DeepSeek-V3.1-Base | DeepSeek | 45.9 | 4 | 2026-06-13 |
| 8 | DeepSeek-V3.2-Exp-Base | DeepSeek | 45.7 | 4 | 2026-06-13 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.