Claw-Eval — Averaged Claw-Eval score across subtasks; distinct axis from pass@3 single-shot metric
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.5 | Anthropic | 76.6 | 1 | 2026-06-15 |
| 2 | Qwen3.6 27B | Alibaba | 72.4 | 1 | 2026-06-15 |
| 3 | Qwen3.5 397B A17B | Alibaba | 70.7 | 1 | 2026-06-15 |
| 4 | Qwen3.6 35B A3B | Alibaba | 68.7 | 2 | 2026-06-15 |
| 5 | Qwen3.5 9B | Alibaba | 66.5 | 2 | 2026-08-10 |
| 6 | Qwen3.5 35B A3B | Alibaba | 65.4 | 1 | 2026-05-03 |
| 7 | Qwen3.5 27B | Alibaba | 64.3 | 2 | 2026-06-15 |
| 8 | LFM2.5-2.6B | Liquid AI | 62.9 | 2 | 2026-08-10 |
| 9 | Qwen3.5 4B | Alibaba | 62.3 | 2 | 2026-08-10 |
| 10 | Gemma 4 26B A4B | 58.8 | 1 | 2026-05-03 | |
| 11 | Gemma 4 E4B | 58 | 2 | 2026-08-10 | |
| 12 | Gemma 4 E2B | 53.1 | 2 | 2026-08-10 | |
| 13 | Gemma 4 31B | 48.5 | 2 | 2026-06-15 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.