GraphWalks — BFS task, 256K-token context subset — Same BFS/256K as 385 (vs 1M); matches graphwalks_bfs_256k, min 73.7 exact.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Mythos 5 | Anthropic | 91.1 | 3 | 2026-06-12 |
| 2 | Claude Opus 4.8 | Anthropic | 85.9 | 3 | 2026-06-12 |
| 3 | Claude Opus 4.7 | Anthropic | 76.9 | 2 | 2026-06-01 |
| 4 | GPT-5.5 | OpenAI | 73.7 | 3 | 2026-06-12 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.