ZebraLogic — Distinct named logic-puzzle benchmark; MiniMax M1/Kimi K2 cards
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Qwen3 VL 235B A22B Reasoning | Alibaba | 97.3 | 1 | 2026-08-23 |
| 2 | o3 | OpenAI | 95.8 | 2 | 2026-06-05 |
| 3 | DeepSeek-R1 | DeepSeek | 95.1 | 4 | 2026-06-05 |
| 4 | Claude 4 Opus | Anthropic | 95.1 | 4 | 2026-06-15 |
| 5 | Qwen3 235B A22B Instruct 2507 | Alibaba | 95 | 1 | 2026-08-23 |
| 6 | Gemini 2.5 Pro Preview 06.05 (32k think) | 91.6 | 2 | 2026-06-05 | |
| 7 | Kimi K2 Instruct | Moonshot | 89 | 4 | 2026-08-23 |
| 8 | MiniMax M1 80K | MiniMax | 86.8 | 3 | 2026-08-23 |
| 9 | DeepSeek-V3 | DeepSeek | 84 | 2 | 2026-06-15 |
| 10 | Qwen3 235B A22B | Alibaba | 80.3 | 4 | 2026-06-15 |
| 11 | MiniMax M1 40K | MiniMax | 80.1 | 3 | 2026-08-23 |
| 12 | Claude Sonnet 4 | Anthropic | 73.7 | 2 | 2026-06-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.