BigCodeBench — Unspecified split (raws fold Pass@1 phrasing); ambiguous vs named Full/Hard subsets, kept apart
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | MiMo V2 Flash Base | Xiaomi | 70.1 | 2 | 2026-06-13 |
| 2 | DeepSeek-V3.2-Base | DeepSeek | 63.9 | 4 | 2026-06-27 |
| 3 | DeepSeek-V3.1-Base | DeepSeek | 63 | 2 | 2026-06-13 |
| 4 | DeepSeek-V3.2-Exp-Base | DeepSeek | 62.9 | 2 | 2026-06-13 |
| 5 | Kimi K2 Base | Moonshot | 61.7 | 2 | 2026-06-13 |
| 6 | DeepSeek-V4-Pro-Base | DeepSeek | 59.2 | 4 | 2026-06-27 |
| 7 | DeepSeek-V4-Flash-Base | DeepSeek | 56.8 | 4 | 2026-06-27 |
| 8 | Granite 4.1 30B | IBM | 38.8 | 1 | 2026-06-15 |
| 9 | Granite 4.1 8B | IBM | 35 | 1 | 2026-06-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.