AutoLogi — Published automated logical-reasoning benchmark
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Sonnet 4 | Anthropic | 89.8 | 2 | 2026-06-15 |
| 2 | Kimi K2 Instruct | Moonshot | 89.5 | 4 | 2026-08-23 |
| 3 | DeepSeek-V3 | DeepSeek | 88.9 | 2 | 2026-06-15 |
| 4 | Claude 4 Opus | Anthropic | 86.1 | 2 | 2026-06-15 |
| 5 | Qwen3 235B A22B | Alibaba | 83.3 | 2 | 2026-06-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.