Legal Agent Benchmark, Harvey's Held-Out Set — Distinct, much harder subset (0.8-13.3) than Full Public Set (8-16.9) - correctly separate.
| # | Model | Vendor | Best score | Runs | Last seen |
|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 13.3 | 1 | 2026-06-10 |
| 2 | Claude Sonnet 4.6 | Anthropic | 5.4 | 1 | 2026-06-30 |
| 3 | GPT-5.5 | OpenAI | 2.1 | 1 | 2026-06-30 |
| 4 | Gemini 3.5 Flash | 0.8 | 1 | 2026-06-30 |
Best tracked score per model (default configuration; source-attributed and verification-tiered). Open a model for its full benchmark surface, provenance and pricing.