FRAMES (Google multi-hop retrieval+reasoning QA) — Same profile (58.1/85/87, 4 models) as 'Frames w/ tools' (351) — likely same run.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Kimi K2 (Reasoning) | Moonshot | 87 | 2 | 2026-05-15 |
| 2 | GPT-5 | OpenAI | 86 | 1 | 2026-05-15 |
| 3 | Claude Sonnet 4.5 | Anthropic | 85 | 1 | 2026-05-15 |
| 4 | DeepSeek-V3.2 | DeepSeek | 80.2 | 1 | 2026-05-15 |
| 5 | DeepSeek-V3 | DeepSeek | 73.3 | 1 | 2026-08-23 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.