CMath — Chinese elementary-school math word-problem benchmark, base metric
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Hunyuan Large | Tencent | 91.3 | 1 | 2026-05-31 |
| 2 | Hy3 preview-Base | Tencent | 91.2 | 2 | 2026-06-05 |
| 3 | Kimi K2 Base | Moonshot | 90.8 | 2 | 2026-06-05 |
| 4 | GLM-4.5-Base | Z.ai | 89.3 | 2 | 2026-06-05 |
| 5 | DeepSeek-V3-Base | DeepSeek | 85.5 | 2 | 2026-06-05 |
| 6 | Moonlight | Moonshot | 81.1 | 2 | 2026-06-12 |
| 7 | Qwen2.5 3B | Alibaba | 80 | 2 | 2026-06-05 |
| 8 | DeepSeek-V2 | DeepSeek | 78.7 | 1 | 2026-05-31 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.