Seal-0 — Search/tool-use benchmark; distinct from the 'w/ tools' condition variant.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Kimi K2.5 | Moonshot | 57.4 | 2 | 2026-08-23 |
| 2 | Kimi K2 (Reasoning) | Moonshot | 56.3 | 2 | 2026-05-15 |
| 3 | Claude Sonnet 4.5 | Anthropic | 53.4 | 1 | 2026-05-15 |
| 4 | GPT-5 | OpenAI | 51.4 | 1 | 2026-05-15 |
| 5 | DeepSeek-V3.2 | DeepSeek | 49.5 | 2 | 2026-06-15 |
| 6 | Claude Opus 4.5 | Anthropic | 47.7 | 1 | 2026-06-15 |
| 7 | Qwen3.5 27B | Alibaba | 47.2 | 1 | 2026-08-23 |
| 8 | Qwen3.5 397B A17B | Alibaba | 46.9 | 1 | 2026-08-23 |
| 9 | Gemini 3 Pro | 45.5 | 1 | 2026-06-15 | |
| 10 | GPT-5.2 | OpenAI | 45 | 1 | 2026-08-24 |
| 11 | Qwen3.5 122B A10B | Alibaba | 44.1 | 1 | 2026-08-23 |
| 12 | Qwen3.5 35B A3B | Alibaba | 41.4 | 1 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.