BigCodeBench (Hard split) — Explicit Hard split, much lower score range confirms harder subset distinct from Full
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen 2.5 Coder 32B Instruct | Alibaba | 27 | 1 | 2026-08-23 |
| 2 | Seed Coder 8B Instruct | ByteDance | 26.4 | 1 | 2026-06-05 |
| 3 | Qwen3 8B | Alibaba | 23 | 1 | 2026-06-05 |
| 4 | Yi Coder 9B Chat | 01.AI | 17.6 | 1 | 2026-06-05 |
| 5 | DeepSeek-Coder-6.7B-Instruct | DeepSeek | 15.5 | 1 | 2026-06-05 |
| 6 | Llama 3.1 8B Instruct | Meta | 13.5 | 1 | 2026-06-05 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.