BIG-Bench Hard — IBM Granite-3.1 spelling of BBH; kept apart from acronym 'bbh' entry (diff population, cautious)
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Llama 3.1 8B Instruct | Meta | 73.4 | 1 | 2026-06-11 |
| 2 | Granite 3.2 8B Instruct | IBM | 71.9 | 1 | 2026-06-11 |
| 3 | Granite 3.1 8B Instruct | IBM | 69.9 | 1 | 2026-06-11 |
| 4 | Granite 3.3 8B Instruct | IBM | 69.1 | 1 | 2026-06-11 |
| 5 | DeepSeek-R1-Distill-Llama-8B | DeepSeek | 67.4 | 1 | 2026-06-11 |
| 6 | DeepSeek-R1-Distill-Qwen-7B | DeepSeek | 67.4 | 1 | 2026-06-11 |
| 7 | Granite 3.3 2B Instruct | IBM | 63.9 | 1 | 2026-06-11 |
| 8 | Granite 3.1 2B Instruct | IBM | 61.8 | 1 | 2026-06-11 |
| 9 | Granite 3.2 2B Instruct | IBM | 61.4 | 1 | 2026-06-11 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.