MT-Bench — LLM-judge multi-turn eval; max=98 is a likely-contaminated outlier vs median=9.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Hunyuan Large Inst. | Tencent | 9.4 | 1 | 2026-05-31 |
| 2 | Llama 3.1 Instruct 405B | Meta | 9.1 | 1 | 2026-05-31 |
| 3 | DeepSeek-V2.5 | DeepSeek | 9 | 2 | 2026-08-23 |
| 4 | Llama 3.1 70B Instruct | Meta | 8.8 | 1 | 2026-05-31 |
| 5 | LFM2.5-8B-A1B-DSpark | Liquid AI | 8.5 | 2 | 2026-08-21 |
| 6 | LFM2.5-2.6B-DSpark | Liquid AI | 5.1 | 2 | 2026-08-21 |
| 7 | LFM2.5-2.6B | Liquid AI | 2.9 | 1 | 2026-08-21 |
| 8 | Qwen2.5 Instruct 72B | Alibaba | 0.9 | 1 | 2026-08-23 |
| 9 | Mistral Large 2 | Mistral | 0.9 | 1 | 2026-08-23 |
| 10 | Qwen2 7B Instruct | Alibaba | 0.8 | 1 | 2026-08-23 |
| 11 | Llama 3.1 Nemotron Nano 8B V1 | NVIDIA | 0.8 | 1 | 2026-08-23 |
| 12 | Llama 3.1 Nemotron Instruct 70B | NVIDIA | 0.1 | 1 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.