GraphWalks — Parents task, 256K-token context subset — Same task+subset as 389 ('256K' vs '256K subset'); ranges overlap (90.1-100 vs 90.1-99.9).
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Mythos 5 | Anthropic | 100 | 3 | 2026-06-12 |
| 2 | Claude Opus 4.8 | Anthropic | 99.3 | 3 | 2026-06-12 |
| 3 | Claude Opus 4.7 | Anthropic | 93.6 | 2 | 2026-06-01 |
| 4 | GPT-5.5 | OpenAI | 90.1 | 3 | 2026-06-12 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.