BabyVision — raws confirm 'BabyVision' capitalized alias; w/-python tools variant kept separate below
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Kimi K3 | Moonshot | 85.7 | 1 | 2026-08-23 |
| 2 | Qwen3.8 27B | Alibaba | 85.6 | 2 | 2026-08-24 |
| 3 | Muse Spark 1.1 | Meta | 76.3 | 1 | 2026-08-23 |
| 4 | muse-glimmer-30b | Meta | 70.4 | 1 | 2026-08-24 |
| 5 | Qwen3.7 Plus Preview | Alibaba | 70.4 | 2 | 2026-08-24 |
| 6 | Kimi K2.6 | Moonshot | 68.5 | 2 | 2026-08-23 |
| 7 | Gemini 3.1 Pro | 51.6 | 2 | 2026-05-03 | |
| 8 | GPT-5.4 | OpenAI | 49.7 | 1 | 2026-05-03 |
| 9 | Qwen3.5 27B | Alibaba | 44.6 | 1 | 2026-08-23 |
| 10 | Qwen3.5 122B A10B | Alibaba | 40.2 | 1 | 2026-08-23 |
| 11 | Qwen3.5 35B A3B | Alibaba | 38.4 | 1 | 2026-08-23 |
| 12 | Kimi K2.5 | Moonshot | 36.5 | 1 | 2026-05-03 |
| 13 | Qwen3.6 27B | Alibaba | 28.9 | 1 | 2026-08-24 |
| 14 | Claude Opus 4.6 | Anthropic | 14.8 | 2 | 2026-08-24 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.