| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.5 | Anthropic | 88.9 | 1 | 2026-07-29 |
| 2 | Claude Opus 4.1 | Anthropic | 86.8 | 1 | 2026-07-29 |
| 3 | Claude Sonnet 4.5 | Anthropic | 86.2 | 1 | 2026-07-29 |
| 4 | Gemini 3 Pro | 85.3 | 1 | 2026-07-29 | |
| 5 | LFM2.5-350M | Liquid AI | 17.8 | 2 | 2026-08-09 |
| 6 | LFM2.5-230M | Liquid AI | 13.7 | 1 | 2026-08-09 |
| 7 | Qwen3.5 0.8B | Alibaba | 7 | 1 | 2026-08-09 |
| 8 | Gemma 3 1B IT | 6.4 | 2 | 2026-08-09 | |
| 9 | Qwen3.5 0.8B (Instruct) | Alibaba | 6.1 | 2 | 2026-08-09 |
| 10 | Granite 4.0 350M | IBM | 6.1 | 2 | 2026-08-09 |
| 11 | LFM2-350M | Liquid AI | 5.6 | 2 | 2026-08-09 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.