GraphWalks — BFS task, 1M-token context subset — Long-context graph-traversal benchmark, BFS task at 1M subset.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Mythos 5 | Anthropic | 74.3 | 1 | 2026-06-12 |
| 2 | Claude Opus 4.8 | Anthropic | 68.1 | 1 | 2026-06-12 |
| 3 | GPT-5.5 | OpenAI | 45.4 | 1 | 2026-06-12 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.