ZeroBench — Kimi K2.5 card; base (no-tool) run
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Kimi K2.5 | Moonshot | 9 | 2 | 2026-08-23 |
| 2 | GPT-5.2 | OpenAI | 9 | 1 | 2026-06-15 |
| 3 | Gemini 3 Pro | 8 | 1 | 2026-06-15 | |
| 4 | Qwen3 VL 235B A22B Reasoning | Alibaba | 4 | 2 | 2026-08-23 |
| 5 | Claude Opus 4.5 | Anthropic | 3 | 1 | 2026-06-15 |
| 6 | Kimi K3 | Moonshot | 0.4 | 1 | 2026-08-23 |
| 7 | DeepSeek-V4-Flash-Vision-Exp | DeepSeek | 0.3 | 1 | 2026-08-23 |
| 8 | Muse Spark | Meta | 0.3 | 1 | 2026-08-23 |
| 9 | Qwen3.5 27B | Alibaba | 0.1 | 1 | 2026-08-23 |
| 10 | Qwen3.5 122B A10B | Alibaba | 0.1 | 1 | 2026-08-23 |
| 11 | Qwen3.5 35B A3B | Alibaba | 0.1 | 1 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.