MMLU-Pro as measured by the HF Open LLM Leaderboard v2 (5-shot direct). Values are RAW accuracy, recovered from the published normalization (raw = norm×0.9+10; baseline 10%, verified). Protocol match to standard MMLU-Pro unconfirmed → unpooled.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen 14B B100 | Alibaba | 51.8 | 1 | 2026-07-04 |
| 2 | Yi 1.5 34B 32K | 01.AI | 47.1 | 1 | 2026-07-04 |
| 3 | Yi 1.5 34B | 01.AI | 46.7 | 1 | 2026-07-04 |
| 4 | Yi 1.5 34B Chat 16K | 01.AI | 45.5 | 1 | 2026-07-04 |
| 5 | Yi 1.5 34B Chat | 01.AI | 45.2 | 1 | 2026-07-04 |
| 6 | Marco-o1 | AIDC-AI | 41.2 | 1 | 2026-07-04 |
| 7 | Yi 1.5 9B Chat 16K | 01.AI | 39.9 | 1 | 2026-07-04 |
| 8 | Yi 1.5 9B Chat | 01.AI | 39.8 | 1 | 2026-07-04 |
| 9 | Yi 1.5 9B | 01.AI | 39.2 | 1 | 2026-07-04 |
| 10 | Qwen2.5 7B Test Novelist | Alibaba | 38.7 | 1 | 2026-07-04 |
| 11 | Yi 1.5 9B 32K | 01.AI | 37.6 | 1 | 2026-07-04 |
| 12 | Yi 9B 200K | 01.AI | 36.2 | 1 | 2026-07-04 |
| 13 | Yi 9B | 01.AI | 35.7 | 1 | 2026-07-04 |
| 14 | smartllama3.1-8B-001 | Meta | 34.9 | 1 | 2026-07-04 |
| 15 | Yi 1.5 6B Chat | 01.AI | 31.9 | 1 | 2026-07-04 |
| 16 | Yi 1.5 6B | 01.AI | 31.4 | 1 | 2026-07-04 |
| 17 | NuminaMath-7B-CoT | AI-MO | 28.7 | 1 | 2026-07-04 |
| 18 | Qwen2.5 1.5B Continuous Learnt | Alibaba | 28.1 | 1 | 2026-07-04 |
| 19 | NuminaMath-7B-TIR | AI-MO | 27.3 | 1 | 2026-07-04 |
| 20 | Llama 3 Instruct 8B | Meta | 26 | 1 | 2026-07-04 |
| 21 | Yi Coder 9B Chat | 01.AI | 24.3 | 1 | 2026-07-04 |
| 22 | Llama Squared 8B | Meta | 23.7 | 1 | 2026-07-04 |
| 23 | Llama 3.1 8B Squareroot | Meta | 17.5 | 1 | 2026-07-04 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.