CLUEWSC — Same CLUEWSC EM metric; DeepSeek internal column name vs display name
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | R1-Zero | DeepSeek | 93.1 | 1 | 2026-05-03 |
| 2 | DeepSeek-R1 | DeepSeek | 92.8 | 6 | 2026-06-04 |
| 3 | Qwen2.5 Instruct 72B | Alibaba | 91.4 | 2 | 2026-05-03 |
| 4 | DeepSeek-V3 | DeepSeek | 90.9 | 8 | 2026-08-23 |
| 5 | DeepSeek-V2.5 | DeepSeek | 90.4 | 1 | 2026-05-03 |
| 6 | OpenAI o1-mini | OpenAI | 89.9 | 5 | 2026-06-04 |
| 7 | DeepSeek-V2 | DeepSeek | 89.9 | 2 | 2026-05-03 |
| 8 | GPT-4o | OpenAI | 87.9 | 6 | 2026-06-04 |
| 9 | Claude 3.5 Sonnet | Anthropic | 85.4 | 6 | 2026-06-04 |
| 10 | DeepSeek-V4-Pro-Base | DeepSeek | 85.2 | 4 | 2026-06-27 |
| 11 | Llama 3.1 Instruct 405B | Meta | 84.7 | 2 | 2026-05-03 |
| 12 | DeepSeek-V3.2-Base | DeepSeek | 83.5 | 4 | 2026-06-27 |
| 13 | DeepSeek-V4-Flash-Base | DeepSeek | 82.2 | 4 | 2026-06-27 |
| 14 | ERNIE 4.5 | Baidu | 48.6 | 1 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.